Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:07:58.092999Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.09233.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:07:58.092999Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 78c7cf6f-a6fd-41d7-8194-20740a501d0c · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models On-policy distillation of language models: Learning from self- generated mistakes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8cd61ef-1c0a-4b01-9dcf-d839d166546d · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models This assumption holds when the reference retains the generative structure of the teacher
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a0d031f-33b6-42e3-b8ff-933126df6d5c · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Minillm: Knowledge distillation of large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fbba7421-c207-4c4f-97c5-1d4ade639580 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Clipscore: A reference-free evaluation metric for image captioning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac58b56-95c9-48d7-a367-578bbfef2d61 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0817feb4-3071-4dc9-8bbd-977b1c309597 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 293641e2-7c0c-48a2-8871-91de8859ea25 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models A Survey of On-Policy Distillation for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9cb6daf-0c3b-4648-ac25-46ebae476dfd · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Score-Based Generative Modeling through Stochastic Differential Equations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dad06f8-b728-49c1-8280-26c03793b422 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Human preference score: Better aligning text-to-image models with human preference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7e1778-c493-4f70-a9d9-1655f7239038 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Advantage weighted matching: Aligning rl with pretraining in diffusion models.arXiv preprint arXiv:2509.25050, 2025a
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a272bc2d-b06b-4b55-b4e1-ca33a95b1cac · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models DiffusionNFT: Online Diffusion Reinforcement with Forward Process
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ad649c-96c9-4f6f-8a31-d5b648fd47bc · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models TheAesthetics teacheris also trained using GRPO-Guard and optimizes the equally weighted reward of PickScore, ClipScore and HPSv2.1
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 47aecb0a-8794-4644-a0ae-ae05181012a8 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models 21 Preprint
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad03b2e0-390d-4837-b2b3-c8881a7d2709 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models It is also used only for out-of-domain evaluation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 69ec120a-db5e-4fa7-8f01-ec5cffe7bb36 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models HELLO" pinned to a denim jacket, close-up macro shot. A cinematic neon-lit ramen shop at night in the rain, a glowing sign reading
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dbbf3af3-0c70-4598-9076-c6df51f939ce · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models MARBLE: Multi-Aspect Reward Balance for Diffusion RL
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a3594f4d-9a00-4f05-8f63-865131c91d43 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3312d3c-18f1-4985-b7f4-74cc0859b3b2 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Flow-OPD: On-Policy Distillation for Flow Matching Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f4f055b-6e15-4e82-b09f-3820823d45d2 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Training diffusion models with reinforcement learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 612c2afb-acef-4ebf-8aa0-b362f70599ee · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Directly fine-tuning diffusion models on differentiable rewards
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55fb6dab-3981-4d66-bef1-c4cc651d7052 · outbound
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models Aligning Text-to-Image Diffusion Models with Reward Backpropagation
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.