Pith. sign in

Paper Citation Record · LEDGER

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

As of 29 July 2026, this Paper Citation Record lists 31 of 31 outbound references and 16 inbound Pith citation observations for arXiv:2604.08527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.08527 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:27:37.161657Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T11:44:54.717393Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:09:56.995110Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact24
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b8abbb58-5eb4-4361-9904-10ebfe911eac · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Process Reinforcement through Implicit Rewards

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:31.624298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:23a32f2fbb46bf8a9eb87d7af550071f2c41e4beb87938d63365822c74ce6491

Observation c04e973e-a9d1-491a-ac75-c8ed7a09058b · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models MiniLLM: On-Policy Distillation of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:40:27.990750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:6a61797417be985cb4a358c22acc469cbd97d1ce4b775cf72c52ffab2f4e8c6b

Observation 2276962d-227c-4e08-842e-ec06895aab23 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:25.498353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:cf3362937c5607ba07c0f2525aafd1a226ddac2ea5a4fa4017d56e2ff9c3f0d4

Observation 81e6947b-18df-4a56-a94d-6cd63946a289 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:38:21.105975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:6949f5c6d2ddd20e9d6d1ed2c8ada49ff4ccb12044ce3c37749c463e292f16b3

Observation 40ac80b8-9e99-4fe2-986d-fe1ded2cef6d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:27.172350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:3857a5008bff7113218bcc8ee8bcda0bccb7940d79ad30d9b50e848d9c5f6eb5

Observation 59c794f9-6257-4678-b11f-b01ad56c2cf5 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Distilling the Knowledge in a Neural Network

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:25.096749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:0b4dd543b23883c1bf66376f4de9ee1eb27f08a0b30ee9eefec8acb8b12c6430

Observation bd51b5a3-88d4-451f-849e-2be9eec2343c · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:59:02.400889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:ea6cfea9f8d9a2b011f84e57b6efd7e9f5351161380432f75d1be70caa0fe5aa

Observation e9156fda-705c-4af2-8f32-7511be2ed81c · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Reinforcement Learning via Self-Distillation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:6a9f4af9596aa2438d5185cff02f6a73cd1fe1729ca7d5896db440ff1c53ddd6

Observation c1fcf864-8175-41a4-87a5-4721055c6e48 · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:25.667356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:33c1a41794cf149e8dc9742560d28c843251aedd276ebc4013201bcc3aae6689

Observation fc0f537e-63d8-4d7b-a27e-e65b487925ae · outbound

This paper cites and Rush, A.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models and Rush, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T10:21:45.409000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:cf26f562749547da73a4a3b5d09db5c98d9dccc723c09ecffab67cd8c34d3fd5

Observation 881ca87b-091f-45a8-ab42-cd3a3ff62932 · outbound

This paper cites Dual Policy Distillation.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Dual Policy Distillation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:26.062388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:105b0086e6373f1f3be4c617a342ba9567e9300cd21d9f58feede9a109d6c1b6

Observation 8181b803-58df-495b-bc2e-104fe803fa31 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Solving Quantitative Reasoning Problems with Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T22:43:59.458883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:5e94e5cc4d42de27bf69c708f4cbca091afa3a1ea3d3fbf40fb75121125e2129

Observation 881f9d32-a398-4993-b33f-c78f028855dd · outbound

This paper cites Let's Verify Step by Step.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Let's Verify Step by Step

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:27.346609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:80485ba3bed2cff5e98ee5d36de119a7b8d1baa4cc4408b8e75387019f011c43

Observation 2970a45c-60c2-4bf4-b696-e486ed827aa7 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:24.521350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:bad295cfdf770f9703b768a0013c6a4dceca1139f1f2175d77446de799cb9f0c

Observation 1c730f72-9819-4b20-b8a5-faa1acdf2f9e · outbound

This paper cites 20251026.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models 20251026

Reference 15

Resolution
malformed identifier
doi_truncated, observed 2026-05-10T17:30:39.818743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:61fd6e277808cc2140c169637e663f23c54ead59893d25ed7aca27c2926dddbd

Observation 9cc87e02-f16b-4500-b8e8-8fe9f02b7009 · outbound

This paper cites Policy Distillation.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Policy Distillation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:28.044908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:034c7c71cf0b67ba2a9d2d77f149957fc3e8666cff5f19a0ed95e4339c6969ce

Observation eb0047b6-7d34-41bb-8d84-d375a8a67b2c · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:24.677843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:9f46517e2f2aede24e56de7aa1d3218983dce8af9e0ce1ba02e3974083773ddf

Observation ca188012-58ae-4272-8959-49e766c006ec · outbound

This paper cites Proximal Policy Optimization Algorithms.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Proximal Policy Optimization Algorithms

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:28.199747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:d2da8ec4a00ee8c5c48a1f37371f9345cfdc1461c157aa6c08cb98d00e4669b2

Observation b5f0a30a-7d66-4bde-9199-c3c70396ac03 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:25.326349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:43af1fbe7a4f836a1a803f30998250edb9d038175dbefa4b0487e4f30fcf8adf

Observation 6dc57683-d61d-453f-a22a-1cc743f085d3 · outbound

This paper cites MiMo-V2-Flash Technical Report.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models MiMo-V2-Flash Technical Report

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:33:32.947936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:3f8d5a3478fa74f5f11b3c61e4190a6bd5d50298525d2337884fd6cd9760664b

Observation 2be2d9c2-6261-4c1e-bdfe-9d85220e63c8 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Learning to Reason under Off-Policy Guidance

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:17:03.075876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:14a37e20d9fdf6f702753011a6fb210fb33043a25fdc91968feac5411d199ffd

Observation 76a7a0f5-5d7e-4777-b998-045d16356060 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:15:55.036587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:928da0f43632837ed2a61ad0513f8b5165ac8755a0ea8cfbb0a0b5af985a8e1c

Observation c134a4ce-be53-4262-8b9f-e6cdc509d312 · outbound

This paper cites Qwen3 Technical Report.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Qwen3 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:28.180751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:15033051e3a7794bf196b77296d78a15b02526bc67e66c6b92b2a31e2bc81b49

Observation 702bf885-7308-45e7-b79b-bde35918684c · outbound

This paper cites Black-box on-policy distillation of large language models.arXiv preprint.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Black-box on-policy distillation of large language models.arXiv preprint

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:26.859445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:64639adfa9a4e2f7341a3a68a5b11641c6b742cf7fa1a1baedd3c27625d422f5

Observation deb75af3-a730-4bbb-910a-c7712e56e9db · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:46:27.522343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:ccaabe19086d7930a8758cf2664943e80c3e2beecca730b2d416980ee7308632

Observation 930f1c66-095d-4c67-a699-e775291e48e7 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:21:05.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:4ccbde3774009fd4416ad046af86ac389d78b04fa5b79633e223b418cc31cd95

Observation 94c4fcd5-d4e2-4e3e-b450-c0537958cc2c · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models A Survey of Reinforcement Learning for Large Reasoning Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.511734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:e193e47d655bb4ef93c25a61e1d135ab57ff7c3e33d3e08ffb0f04bf59d47779

Observation a6f8d663-b49c-4d2f-8e24-bc00a140300f · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:54:31.187112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:c3f9e95974fa9c33e60d981a03fd55f04f1854309b545cbf2fced3e2d4a630cd

Observation 936e03b1-0802-4aba-8450-526f25379efd · outbound

This paper cites Appendix B.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Appendix B

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T10:21:45.417635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:4bec499586a0490fe870acb2b0d1438a6980c12a363dbb1b07155340c4b13830

Observation 36e8de24-1349-4f57-8d75-c720b8018423 · outbound

This paper cites an unresolved cited work.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Unresolved cited work

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T10:21:45.411892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:ad2f32fa6c3473953bb00ca08b4b0365d60578e1daa8f35d2810193e1127278e

Observation f8b954ce-30bb-46e4-bf29-10df51c99bcf · outbound

This paper cites an unresolved cited work.

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Unresolved cited work

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T10:21:45.414892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-10T17:27:37.161657Z digest=sha256:ec9c35103b1ebad836de3d8d3406c335427f855ccdd31e1ecbf45b9ae2087032

Pith citing papers

Observation bffb3a39-fc4f-4bc8-aa3d-0a6d1dc7315e · inbound

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models cites this paper.

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:11:12.902477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-08T09:57:17.662754Z digest=sha256:fc2403f672d27c7a49a85c1190a73f8b7d6c61077d5fc17ca137b33b9f53cdf3

Observation 5789fdb2-a335-4cac-a115-a4d4bcef9fb3 · inbound

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models cites this paper.

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:46:22.544088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T09:44:55.200796Z digest=sha256:d4191fe1be30a22414776432699d0e3599fe4c30a44af8a55b71c49b4859944a

Observation 9d78977c-c3e9-4d74-a39f-68247d4e9259 · inbound

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes cites this paper.

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:22:06.953324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T02:18:29.486408Z digest=sha256:bf28802e7880a16693a5a0fda8f463456e5b42faecdcc366544431c3bad99cb1

Observation f1800a1b-5e5e-485b-a988-6c6fbb5b4e6c · inbound

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes cites this paper.

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.400937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-30T22:14:50.636752Z digest=sha256:de91eaaaa6de0374ca9466a9ca9e4dbdf41cc91ee00ae4b165ffe4d8c373b0ea

Observation 8e3ab833-e9bd-4aca-809f-72e9b236ea08 · inbound

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation cites this paper.

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:52:46.201324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-19T20:49:15.724896Z digest=sha256:ce912232c2dd4e0562e7fb2b757fe9a240af682def0d90464fecf4e486d4fb4a

Observation 93ef7496-5ceb-4dbe-b906-15219d6078a4 · inbound

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification cites this paper.

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:12:24.145017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:a6da0f27a196701b4a6d629e663d8e02829f5a5cb8f6c3c7f54a7df3f86b25e0

Observation d74c0cfe-4275-404a-b72d-b72efe45aa59 · inbound

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation cites this paper.

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:17.360936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-06-28T15:14:42.489647Z digest=sha256:a1642accc830907baa47531764fbac74009bfe67babeb2cff50907ab8ee506a0

Observation 0047499b-cde7-48d1-8bcc-2310a655b21d · inbound

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation cites this paper.

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:36:57.003582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-28T01:59:15.154873Z digest=sha256:4adc7f3d7e10ea1addc461a389fba02ccc6abd4133e1d02dddbe5ac1b8b76689

Observation 9d842b35-1215-49f0-af28-22c9e37cf14e · inbound

SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation cites this paper.

SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-06-26T20:29:57.859170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:d8f3d7efc64f39c660f630efb29145166e39ad2f6aa29f1b0b8863dd4ad2223c

Observation 9a7e96b4-03d0-4a8d-bf88-d2c65942ea5b · inbound

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning cites this paper.

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:39:42.400353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T11:08:54.744334Z digest=sha256:e4ee01e3abeedaf4404d99daf534faf4c4d623333a51ad4f4fb394209ea3a83a

Observation 4eef3ad0-b655-400e-9bbd-3ac125455005 · inbound

A Formula-Driven Survey and Research Agenda for On-Policy Distillation cites this paper.

A Formula-Driven Survey and Research Agenda for On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:19:47.147150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-06-26T09:02:13.340365Z digest=sha256:d4eda4bf470890e62a4076453b8fe646d26d0f7673a4c55c4cd01cd36f6921ae

Observation 7e5ca99a-2617-4633-93af-f0db5377a306 · inbound

Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation cites this paper.

Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.957887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T08:49:58.735297Z digest=sha256:7222cb52571dc435df92236d5bd4c612adaaf5c7b103598d6360376d58cf88fb

Observation be204bb8-3062-423f-84b3-3ce60c6f1b02 · inbound

Blockwise Policy-Drift Gating for On-Policy Distillation cites this paper.

Blockwise Policy-Drift Gating for On-Policy Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:09:56.996729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T00:57:05.042921Z digest=sha256:d73f026c5271b228cef262ee84cb03ccd969c1af39fe4dce69fcac3946a7bf87

Observation 8501e135-1f97-47cf-b856-bfe3bf3fd30f · inbound

DanceOPD: On-Policy Generative Field Distillation cites this paper.

DanceOPD: On-Policy Generative Field Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:49:51.441439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-26T04:55:42.018348Z digest=sha256:6248de36f7debbb4ddfed837de3bace2d24343d73c172980fa5a72e798641c2b

Observation 0eb77a78-3ce9-4d3a-a886-f0c4d2bc17b8 · inbound

DanceOPD: On-Policy Generative Field Distillation cites this paper.

DanceOPD: On-Policy Generative Field Distillation Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T11:44:54.717393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:44:54.717393Z digest=sha256:7de594fa6a352d153a1de5f0be1a23828d9550853af0791c116afe80571fffc5

Observation c825d042-3a3b-469e-a29b-82aeaefddc21 · inbound

Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts cites this paper.

Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:24:27.013111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-06-30T06:05:47.948367Z digest=sha256:909d27598491d87ab768b014dc8e98eea1c106934b383f201aa14d6d5a8cfccc