DreOPD derives a closed-form velocity regression target v* = v_T + (lambda-1)(v_T - v_ref) that lets flow-matching students extrapolate beyond their teachers, with a mildly degraded reference amplifying the direction.
Minillm: Knowledge distillation of large language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models
DreOPD derives a closed-form velocity regression target v* = v_T + (lambda-1)(v_T - v_ref) that lets flow-matching students extrapolate beyond their teachers, with a mildly degraded reference amplifying the direction.