Pith. sign in

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it
abstract

Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to single-task optimization. Extending RL to multiple tasks is challenging: joint optimization suffers from cross-task interference and imbalance, while cascade RL is cumbersome and prone to catastrophic forgetting. We propose DiffusionOPD, a new multi-task training paradigm for diffusion models based on Online Policy Distillation (OPD). DiffusionOPD first trains task-specific teachers independently, then distills their capabilities into a unified student along the student own rollout trajectories. This decouples single-task exploration from multi-task integration and avoids the optimization burden of solving all tasks jointly from scratch. Theoretically, we lift the OPD framework from discrete tokens to continuous-state Markov processes, deriving a closed-form per-step KL objective that unifies both stochastic SDE and deterministic ODE refinement via mean-matching. We formally and empirically demonstrate that this analytic gradient provides lower variance and better generality compared to conventional PPO-style policy gradients. Extensive experiments show that DiffusionOPD consistently surpasses both multi-reward RL and cascade RL baselines in training efficiency and final performance, while achieving state-of-the-art results on all evaluated benchmarks.

fields

cs.CV 3

years

2026 3

representative citing papers

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

cs.CV · 2026-06-16 · unverdicted · novelty 5.0

MaineCoon is presented as the first 22B-parameter real-time streaming audio-visual autoregressive model optimized for social-interactive applications, using novel training techniques and an agentic inference framework.

Qwen-Image-Flash: Beyond Objective Design

cs.CV · 2026-06-02 · unverdicted · novelty 4.0

Empirical analysis of data, guidance, and task mixture in few-step distillation of Qwen-Image-2.0 produces the Qwen-Image-Flash model with improved performance in unified generation and editing tasks.

citing papers explorer

Showing 3 of 3 citing papers.

  • MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model cs.CV · 2026-06-16 · unverdicted · none · ref 26 · internal anchor

    MaineCoon is presented as the first 22B-parameter real-time streaming audio-visual autoregressive model optimized for social-interactive applications, using novel training techniques and an agentic inference framework.

  • Qwen-Image-Flash: Beyond Objective Design cs.CV · 2026-06-02 · unverdicted · none · ref 5 · internal anchor

    Empirical analysis of data, guidance, and task mixture in few-step distillation of Qwen-Image-2.0 produces the Qwen-Image-Flash model with improved performance in unified generation and editing tasks.

  • DanceOPD: On-Policy Generative Field Distillation cs.CV · 2026-06-25 · unreviewed · ref 53 · internal anchor