Pith. sign in

Paper Citation Record · LEDGER

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.09233.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09233 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:07:58.092999Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78c7cf6f-a6fd-41d7-8194-20740a501d0c · outbound

This paper cites On-policy distillation of language models: Learning from self- generated mistakes.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models On-policy distillation of language models: Learning from self- generated mistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.796547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.796547Z digest=sha256:1532c6dbe2062e116a783ac69feb9b9af8fe5c25ce892e6eac6ca6e90b078a1d

Observation e8cd61ef-1c0a-4b01-9dcf-d839d166546d · outbound

This paper cites This assumption holds when the reference retains the generative structure of the teacher.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models This assumption holds when the reference retains the generative structure of the teacher

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:07:58.496244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:07:58.088322Z digest=sha256:8fa4acd251b1087dc5185920782d1cf6f6c858887c9aabb4ac195571dcfcdf99

Observation 0a0d031f-33b6-42e3-b8ff-933126df6d5c · outbound

This paper cites Minillm: Knowledge distillation of large language models.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Minillm: Knowledge distillation of large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:07:58.593182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:07:57.816973Z digest=sha256:5591067a10cc4e46e6f022e4e7eec5fab83bf95db97ec2ea483bd2f60b722043

Observation fbba7421-c207-4c4f-97c5-1d4ade639580 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Clipscore: A reference-free evaluation metric for image captioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.821683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.821683Z digest=sha256:251505fe978e01ef177925313bc65394633f53bc4dc4c0f3bfff543774910e9e

Observation 6ac58b56-95c9-48d7-a367-578bbfef2d61 · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.826018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.826018Z digest=sha256:7cdc63c3ff2bf5fb9ca03c80af34a200462e13a2c286dc5cd5baa390ee1bc781

Observation 0817feb4-3071-4dc9-8bbd-977b1c309597 · outbound

This paper cites DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.830794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.830794Z digest=sha256:91534c212fe2ca14195685350ce8d28af445813eb2892a42ab9e8f5271c161c7

Observation 293641e2-7c0c-48a2-8871-91de8859ea25 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models A Survey of On-Policy Distillation for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.845214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.845214Z digest=sha256:d52aa8515dfcdbc1236396e85647bc89ea7ae5d9767cefd63eb42cdc1d488c9b

Observation e9cb6daf-0c3b-4648-ac25-46ebae476dfd · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Score-Based Generative Modeling through Stochastic Differential Equations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.851263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.851263Z digest=sha256:103f5e27f8f0c9fd336f5290cf2363caf007266b4f47fa3223775d3fc68e7d46

Observation 9dad06f8-b728-49c1-8280-26c03793b422 · outbound

This paper cites Human preference score: Better aligning text-to-image models with human preference.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Human preference score: Better aligning text-to-image models with human preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.933710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.933710Z digest=sha256:327d816523916a6dab9a0628d01d07984319ad60256ab2ce1a85441039acebd1

Observation 0d7e1778-c493-4f70-a9d9-1655f7239038 · outbound

This paper cites Advantage weighted matching: Aligning rl with pretraining in diffusion models.arXiv preprint arXiv:2509.25050, 2025a.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Advantage weighted matching: Aligning rl with pretraining in diffusion models.arXiv preprint arXiv:2509.25050, 2025a

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:58.046455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:58.046455Z digest=sha256:27fce11d5e886234841f78aa3ec8d807dd91cd976fb3ed8e4a84afd46cd30edf

Observation a272bc2d-b06b-4b55-b4e1-ca33a95b1cac · outbound

This paper cites DiffusionNFT: Online Diffusion Reinforcement with Forward Process.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:58.062792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:58.062792Z digest=sha256:523fc049c23fd03a5622ccaf3cb8ea7c2e2e010ab52e23a1d323cf8ce755459c

Observation 96ad649c-96c9-4f6f-8a31-d5b648fd47bc · outbound

This paper cites TheAesthetics teacheris also trained using GRPO-Guard and optimizes the equally weighted reward of PickScore, ClipScore and HPSv2.1.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models TheAesthetics teacheris also trained using GRPO-Guard and optimizes the equally weighted reward of PickScore, ClipScore and HPSv2.1

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:07:58.554633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:07:58.069676Z digest=sha256:2a75c25286b1c4367f36c2a453db22703656b86fa2c9e62fa277e78b91fd8fdf

Observation 47aecb0a-8794-4644-a0ae-ae05181012a8 · outbound

This paper cites 21 Preprint.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models 21 Preprint

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:07:58.539166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:07:58.076359Z digest=sha256:e7afb6852c34e1a7ca1a0a1c41f29ee666325dce989cd37048b411238de86d0c

Observation ad03b2e0-390d-4837-b2b3-c8881a7d2709 · outbound

This paper cites It is also used only for out-of-domain evaluation.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models It is also used only for out-of-domain evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:07:58.523040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:07:58.081865Z digest=sha256:992fe9724928c079abc11f555c0ee873d439d3d93563ba1c4bccebd5d67899af

Observation 69ec120a-db5e-4fa7-8f01-ec5cffe7bb36 · outbound

This paper cites HELLO" pinned to a denim jacket, close-up macro shot. A cinematic neon-lit ramen shop at night in the rain, a glowing sign reading.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models HELLO" pinned to a denim jacket, close-up macro shot. A cinematic neon-lit ramen shop at night in the rain, a glowing sign reading

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:07:58.475946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:07:58.092999Z digest=sha256:703e7c036351f3a1b236f55a5d1d96e3326e38607e69d76c49d4f8ea38f0c384

Observation dbbf3af3-0c70-4598-9076-c6df51f939ce · outbound

This paper cites MARBLE: Multi-Aspect Reward Balance for Diffusion RL.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models MARBLE: Multi-Aspect Reward Balance for Diffusion RL

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:07:58.194038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:07:58.054415Z digest=sha256:da6ec26dccc90dd41d2c9ace89a949ff4cc8da785f7987a639e813577ebb6ae1

Observation a3594f4d-9a00-4f05-8f63-865131c91d43 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.840439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.840439Z digest=sha256:c69dfdfde71504f436047a36fadc7a93d10ad3e56561bf209bd0a7b5a008b35a

Observation a3312d3c-18f1-4985-b7f4-74cc0859b3b2 · outbound

This paper cites Flow-OPD: On-Policy Distillation for Flow Matching Models.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Flow-OPD: On-Policy Distillation for Flow Matching Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.812113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.812113Z digest=sha256:8b31b0961c1890d4af9afbefb8e7c3d253e991bd315125ff69e888976d99c0a6

Observation 6f4f055b-6e15-4e82-b09f-3820823d45d2 · outbound

This paper cites Training diffusion models with reinforcement learning.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Training diffusion models with reinforcement learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.802134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.802134Z digest=sha256:3711fca64f4af5a7b1609c1eadf9ad35cb3d896c5ebc0ee4214209a6b2b3df4b

Observation 612c2afb-acef-4ebf-8aa0-b362f70599ee · outbound

This paper cites Directly fine-tuning diffusion models on differentiable rewards.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Directly fine-tuning diffusion models on differentiable rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.806541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.806541Z digest=sha256:07ee89ad136fdb4d9412d29f3baa11276c596c6a06946bb4ebdc9e7fe5bf0f8f

Observation 55fb6dab-3981-4d66-bef1-c4cc651d7052 · outbound

This paper cites Aligning Text-to-Image Diffusion Models with Reward Backpropagation.

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T21:07:57.835657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:07:57.835657Z digest=sha256:855570a7d7dde508f326cd0cb4f5f5c4d29540c0bf69ef60b5dc5708fac7485a

Pith citing papers

No inbound Pith citation observations are available.