Pith. sign in

Paper Citation Record · LEDGER

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

As of 17 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2607.28449.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28449 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T07:02:36.752153Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddd7fc9e-a4d0-4546-a811-6f358c73364b · outbound

This paper cites AIME 2024.https://huggingface.co/datasets/AI-MO/aimo-validation-aime,.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models AIME 2024.https://huggingface.co/datasets/AI-MO/aimo-validation-aime,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.180458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.180458Z digest=sha256:a3dad75c72428cfb435e285c99dab59ff08ba40ae3c481cee11c632957da9cdc

Observation 0bd79266-7d44-432b-ad48-0a6dd58fc69f · outbound

This paper cites Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.204746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.204746Z digest=sha256:a2653ce7fb305885e018530ff51e3872b16081d6abb899878ffffd43ed6ce033

Observation cc5aab60-667a-4dd0-b780-aea286a13b4c · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models OpenThoughts: Data Recipes for Reasoning Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.269245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.269245Z digest=sha256:abb144db2bd69eabe77ae777c693c9ba8cf04396eb58147a281ffddd87221242

Observation 2cc1a036-9424-4bbf-b503-033469dd48cb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.281808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.281808Z digest=sha256:0812d094e04738401e90ca24bcc85c36488d38ac756823d017b4a406c690215a

Observation 5909f497-0756-4b3b-ad21-d3661e934ae0 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.299150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.299150Z digest=sha256:8efbd320f3943dfb83dc96bab48492b186c6983e2ec2ea1178cd74f95c7dd15d

Observation f0d5da68-1923-4559-a563-0905637bbfdc · outbound

This paper cites Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.307359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.307359Z digest=sha256:cb0d29fffe0e20e128cabeaa782ecaf35b266b2ed0f121579564ac261d9683f6

Observation 975bbc5e-7d34-4bef-8d83-7775c76450b1 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Reinforcement Learning via Self-Distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.316638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.316638Z digest=sha256:273a4b7c29427bdf80d5ebeca1bec965fc0676e0bf3abd3c59f607f510851ff5

Observation 2084da83-caa3-4c49-941a-7d1e56ac3087 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.323569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.323569Z digest=sha256:8ef363f9d3ddd54e9e24a435c9833e973a4f5c9b607ce375249b829c38ea3263

Observation d38dd32c-7209-41ae-af75-6a61e6b03265 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Entropy-Aware On-Policy Distillation of Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.329502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.329502Z digest=sha256:950a933737c1cf1e2222d1868688774479c8da1b8c0eb185dcf50c10da9ff96a

Observation ee7c5b74-6577-4cb7-87cc-c9ea7e5aa1ab · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.338599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.338599Z digest=sha256:003f9463c2df9b0367a24c7495f772d5aea758225f3a2ccf29cd16864084e8ec

Observation 42a823f1-4c4c-42f4-a412-dc775f6283bb · outbound

This paper cites Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.350495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.350495Z digest=sha256:ddd1a6f174496c6fab4890586d8c0956254020a44ab70bc34a6a987845307de8

Observation 45b670b4-ea02-4b7d-afb4-9d72cc8f7a99 · outbound

This paper cites Sequence-level knowledge distillation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Sequence-level knowledge distillation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.358142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.358142Z digest=sha256:8e997c34045f990a46cb739aa0733f872a8afcfbadfa27e3cc190e63fa4617b0

Observation 7862e934-f55d-4afc-820c-8cc32614c136 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.364735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.364735Z digest=sha256:eeaae7fbb2552971b4a8a9bb2e516df4639b0823fe611fa050cf905168d4224b

Observation 807dbbc4-4eec-4054-bdb4-cf5469ab7aff · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.374494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.374494Z digest=sha256:53d89644d79568744e198648c909769acf3de43d5fb7c5dcb1b92b329739deb0

Observation 03ae4232-676a-4c37-96ab-9431c39276a7 · outbound

This paper cites ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.381009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.381009Z digest=sha256:e82b6269541a751c05f9c060dca27198e26624d517b8fbfa881c640a4da1f28b

Observation bbb182fe-34e0-412a-a434-eec9270291e4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.397596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.397596Z digest=sha256:ba9b34eb313c22358a8108dbbed7a71293f02f762743e049e88a5b5b152ef269

Observation 3ff7c33b-6005-4a1b-b8ee-c1adfba11976 · outbound

This paper cites UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models UFT: Unifying supervised and reinforcement fine-tuning.arXiv preprint arXiv:2505.16984, 2025a

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.405842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.405842Z digest=sha256:20187ff6559218f5a765122da3d85cc7562be48446e45a5abeb300fd9831fdcd

Observation b203fea0-f5bf-4942-a61e-1442a19c46d1 · outbound

This paper cites https://thinkingmachines.ai/blog/on-policy-distillation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models https://thinkingmachines.ai/blog/on-policy-distillation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.413600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.413600Z digest=sha256:d7f4fcc9fafc98cf207afb1e1341e6c46f5afc88c923758ec6af9a86a2fd3fd2

Observation 25a087ab-2523-4016-a877-a14e7cf84415 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.435051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.435051Z digest=sha256:e8c890dd7e2e5515f999c85bf70e5f928a752801a824527e50c5829af7d59686

Observation ae4651e4-33f5-43b2-9cd1-451b62bcc1d0 · outbound

This paper cites NVIDIA Nemotron 3: Efficient and Open Intelligence.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models NVIDIA Nemotron 3: Efficient and Open Intelligence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.447129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.447129Z digest=sha256:b3dcc8da92ae4ba38f0f78e8c1c3b41659ed04f71c8bc1a1bddc8c6e2d2f585b

Observation f19ca90c-1a11-44d1-8095-362418d35de6 · outbound

This paper cites AIME 2025.https://github.com/open-compass/opencompass,.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models AIME 2025.https://github.com/open-compass/opencompass,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.457390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.457390Z digest=sha256:0cfeb8dda24b6543bb4ab3ee705f0531303bdf0ceff23414956f89c6da731a69

Observation 89a053cc-c1f2-4931-8fd0-2c31e75c0ac3 · outbound

This paper cites Revealing the power of post-training for small language models via knowledge distillation.arXiv preprint arXiv:2509.26497,.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Revealing the power of post-training for small language models via knowledge distillation.arXiv preprint arXiv:2509.26497,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.476557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.476557Z digest=sha256:3d4207cf0e1fa1032bda11ce717144456a5edbfc9ffd14e67098f7f5bc5f08a1

Observation 9cfac691-02da-416c-82de-29742d00f830 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.483077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.483077Z digest=sha256:89fc01c5abec4d33fe61d2dca38786968a51c63cf30f4608baae8b7aa8663e12

Observation 49ee137f-c17f-450c-9951-754fbe620e4d · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.508401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.508401Z digest=sha256:a0a97220f8a27b2d7b536d8f18cf08db9c60dab27e2737d7ecb48994074b4a6a

Observation 8cccce64-d684-4371-acfd-817a38490354 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Self-Distillation Enables Continual Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.515784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.515784Z digest=sha256:1d141d81298602397c2ba741162b72759b0415b07bb054557aaa757945b3d867

Observation deef98bc-7504-47bb-a9d0-3f9df3ab35df · outbound

This paper cites OpenAI GPT-5 System Card.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models OpenAI GPT-5 System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.526619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.526619Z digest=sha256:641fae463434741f3e43525dee92cf9348b07c32653bbde0bf3a9a4a2fbe41e1

Observation a5d7d686-12d0-4d27-a4c7-8346910dee3b · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models A Survey of On-Policy Distillation for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.537240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.537240Z digest=sha256:fd1336da74f262ea272d4c013bdc09b573fb77ba6d8a46a06992c427b58b8078

Observation 7dc31d68-6bdc-4893-bc14-e94e6b49a585 · outbound

This paper cites Self-Supervised On-Policy Distillation for Reasoning Language Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Self-Supervised On-Policy Distillation for Reasoning Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.563876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.563876Z digest=sha256:2033fe2e8b3657f1745cafb2c410edb184fd5796038e14c020beccf5ead9c8ad

Observation 3e67e821-e832-4136-961d-3518d6577db3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.571677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.571677Z digest=sha256:08051c7db2e991438fa1d5e2da990392f301e9679f1bc6448382812c98bd3006

Observation b99a255f-b801-4986-9efa-96857d746678 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Kimi K2.5: Visual Agentic Intelligence

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.581125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.581125Z digest=sha256:537d6d6b2a654e94577375a7c6821616d93be32ae72e1b6561159da7154cae8e

Observation 0d7ae3d3-3481-46ce-831b-d24f80758b7a · outbound

This paper cites Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.592457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.592457Z digest=sha256:aea1c9ee3071e3ac26b3a8530abf44c042c16684b21d4a62823d0344bde3e108

Observation 6c910826-67b5-4b69-8fe3-5f187761628a · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.604808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.604808Z digest=sha256:d827ba152cad2a3e03f525d96a2ca3c3b0c2e83f055194df5e0204772732cfcd

Observation a39875db-6500-42a1-a564-561822931b6d · outbound

This paper cites MiMo-V2-Flash Technical Report.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models MiMo-V2-Flash Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.615441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.615441Z digest=sha256:be5cdedb2431a6da3170be5bed300dc966f487a58341f0d38cf847392dbe5a2c

Observation a7ab9de9-2145-472b-89ed-1c45fef0917f · outbound

This paper cites On the Position Bias of On-Policy Distillation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models On the Position Bias of On-Policy Distillation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.635289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.635289Z digest=sha256:a941c73f5833194d93d9ad67a3822696507cd19666f0ecc26e2df185946273d7

Observation 77ac798f-70e1-4acb-a73f-48793c51da64 · outbound

This paper cites SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.646126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.646126Z digest=sha256:356f663d26169611144efb273b2160b7027ee4fb3a7bb1c6d8b20cb8a9b477f5

Observation afa89616-2045-4e00-a665-2d49fae5c1eb · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Learning to Reason under Off-Policy Guidance

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.654879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.654879Z digest=sha256:b7ef90761bbceb05cb2909088d821a575e49e5b85c2b66d699ec29b0eed2b8ae

Observation 588ac625-7646-4f1a-b16a-6753604ab89f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.667873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.667873Z digest=sha256:af47b365b3602d6578a22f5b759fdf4d686df9716fdebf20d20c37e7678b4aa8

Observation 20c6b745-2daf-498b-9e44-a798cab22edd · outbound

This paper cites Qwen3 Technical Report.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Qwen3 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.677529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.677529Z digest=sha256:ceb4e61e562c4c3194f5752bdd35396d2eff922d283f30d355a53fe91a0af21c

Observation 185ee781-bb66-4b78-bd75-72e87e492e14 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.688543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.688543Z digest=sha256:03e0cbbfd4f83c9ca940f5170b988e8a05f9525270ad563f1d511d81242ed876

Observation 6dbe2b1a-9005-4219-86c6-c9868c18dbe1 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.698238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.698238Z digest=sha256:e68573ef8e9c327542101ce5a4f2b0df471abe7a61af22888b0df692799bb633

Observation eef9deb5-9289-4a84-8567-56f20e255c95 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.706606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.706606Z digest=sha256:6c81a55e5385b03f08898de68c6de6065cf463b57fc7703cb38937ab72a7bf66

Observation 01e5a407-f9da-4655-ae0e-699014a21359 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.717884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.717884Z digest=sha256:6cd324d4e01f519d6b36b424dbf4fc1d22d26bbb685cd7f78a8bbfd389023c3e

Observation cc9009ee-bca9-460e-8e6a-5dcf8a00d7c2 · outbound

This paper cites On-policy RL meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting.arXiv preprint arXiv:2508.11408,.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models On-policy RL meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting.arXiv preprint arXiv:2508.11408,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.725703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.725703Z digest=sha256:4a3f474bdb3606c03a681cb8995d3afec4d20ab0a254149d22076a07a8c0fb9c

Observation 6fb05be9-1614-43c3-8c22-408743b5e0a0 · outbound

This paper cites Reinforcement-aware Knowledge Distillation for LLM Reasoning.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Reinforcement-aware Knowledge Distillation for LLM Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.733843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.733843Z digest=sha256:2f0ceaf6639013145cf8f375d5493ab03dd65f191dcce7132a1ea9b7a7e1f7d3

Observation c92d4674-260a-4c38-99d5-528589902195 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.741652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.741652Z digest=sha256:5534e2e3ad74ce35b5c080523f4b67f126b77f0473fb626b2caf30fc2929ccb1

Observation 911d5221-e152-47a5-a0ad-af83f2ee42b4 · outbound

This paper cites Group Sequence Policy Optimization.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Group Sequence Policy Optimization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.752153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.752153Z digest=sha256:630f7f8d3a216f3d3858f529eac1c3e8bfa37dfed237c877afa2dfa71b59a392

Observation 05df5b25-d605-4c24-803e-01617e1caa9e · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.290410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.290410Z digest=sha256:ebd45692320615c3499cff0749a0d65bb9362cf48163e2ea45e63f42ef88e9c8

Observation 7ee3de70-d338-4fa4-8329-73a1ae80d57c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.494758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.494758Z digest=sha256:b0ea78f61cf01eb4d9ec860328d08544cfa707c5a9153dfebff3e8e7955d4467

Observation fcc3823e-a6b5-41ca-ae76-9a21050b90c0 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Process Reinforcement through Implicit Rewards

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.247026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.247026Z digest=sha256:4aa4d7654c075ca13d642f75df9e81597fd119e4e8d10fa939b1723eae6c2af7

Observation b30ab58f-eed9-4e45-8784-87573d40aa35 · outbound

This paper cites Privileged Information Distillation for Language Models.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Privileged Information Distillation for Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.469715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.469715Z digest=sha256:c7d0231349c1eec4b06612a0b05fef8a2d6c05679e2cf227ea8fa16aa9c6fba8

Observation 41570081-83f2-4aaf-b8ca-184ecb71579d · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.260317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.260317Z digest=sha256:be6e0e977406e55e684940b4932bdf28670d65be4fd098df5a4a50edacfd5f16

Observation 8c070f38-ab33-438f-a32f-01cbf65e2b24 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.170111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.170111Z digest=sha256:26ac970ec90637cbeae6b45a7d62e3c9a7d97af9a94de87871e5c2c44cb91cd0

Observation 4395ab5a-1095-460b-ab01-5d9f6faca5e6 · outbound

This paper cites Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.233742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.233742Z digest=sha256:ee3020badc882736cd76bde9b5dc14e8235a11d58b93241be1937d7f5b54a179

Observation 9884baaf-7536-45d6-9bb8-5a78833016be · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.224134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.224134Z digest=sha256:79aec2878734049437455737869f20e2afb763e6edff7d41ce7aac1577aff370

Pith citing papers

No inbound Pith citation observations are available.