Pith. sign in

REVIEW 18 cited by

DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16381 v3 pith:ZLAFCHZ7 submitted 2023-05-25 cs.LG cs.CV

classification cs.LGcs.CV
keywords modelsfine-tuningrewardtext-to-imagedpokdiffusionfunctionlearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Learning from human feedback has been shown to improve text-to-image models. These techniques first learn a reward function that captures what humans care about in the task and then improve the models based on the learned reward function. Even though relatively simple approaches (e.g., rejection sampling based on reward scores) have been investigated, fine-tuning text-to-image models with the reward function remains challenging. In this work, we propose using online reinforcement learning (RL) to fine-tune text-to-image models. We focus on diffusion models, defining the fine-tuning task as an RL problem, and updating the pre-trained text-to-image diffusion models using policy gradient to maximize the feedback-trained reward. Our approach, coined DPOK, integrates policy optimization with KL regularization. We conduct an analysis of KL regularization for both RL fine-tuning and supervised fine-tuning. In our experiments, we show that DPOK is generally superior to supervised fine-tuning with respect to both image-text alignment and image quality. Our code is available at https://github.com/google-research/google-research/tree/master/dpok.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.

  2. Efficient Controllable Diffusion via Optimal Classifier Guidance

    cs.LG 2025-05 conditional novelty 7.0 of 10

    SLCD provably converges, under no-regret learning and a strong score-estimation assumption, to the KL-regularized optimal distribution using only supervised classification oracles.

  3. OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A knowledge-graph benchmark (OmniPhys, 1,551 prompts, 14 physics knowledge points) and a batch-feedback prompt optimizer (OmniPrompt) improve measured physical consistency of text-to-image models by 0.01–0.03 Joint Sc...

  4. Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    CACFM applies RL to adaptively select critical regions in probability flow ODE trajectories for consistency distillation, yielding SOTA few-step results on FLUX and SDXL.

  5. X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GRPO reinforcement learning applied to a discrete autoregressive image generator with a diffusion decoder improves instruction following, image quality, and long-text rendering in a unified multimodal model.

  6. Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TACA scales cross-modal attention logits by a timestep-dependent temperature to rebalance text and visual tokens, improving T2I-CompBench alignment on FLUX and SD3.5.

  7. Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    HALO aligns text-to-video diffusion models by jointly optimizing patch-level and video-level DPO rewards, with modest and partially self-referential benchmark gains.

  8. Action-based image editing guided by human instructions

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EditAction fine-tunes InstructPix2Pix with a contrastive action loss and video-derived before/after frames to edit images according to action text commands while preserving object appearance and background.

  9. Bridging the Gap: Aligning Text-to-Image Diffusion Models with Specific Feedback

    cs.CV 2024-11 reject novelty 6.0 of 10

    An object-detection reward that checks category and count fidelity is used to fine-tune Stable Diffusion, but the headline evaluation metric is the same as the training reward.

  10. Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Two plug-and-play strategies — per-timestep advantage weighting and advantage-based trajectory replay — improve diffusion RLHF sample efficiency up to 6× across five reward functions.

  11. Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework

    cs.HC 2025-08 unverdicted novelty 5.0 of 10

    A three-layer framework (input, processing, output) for adaptive external human-machine interfaces in autonomous vehicles is introduced to systematize design and analysis.

  12. ImageReFL: Balancing Quality and Diversity in Human-Aligned Diffusion Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ImageReFL combines base-model early diffusion steps with a real-image-based fine-tuning objective to improve the quality-diversity trade-off in reward-aligned text-to-image generation.

  13. $I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion

    cs.CL 2025-05 reject novelty 5.0 of 10

    A pairwise-conditioned diffusion model generates instructional illustrations from procedural text and is finetuned with a text-image alignment reward.

  14. Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A continuous-time RL algorithm that treats diffusion scores as actions fine-tunes text-to-image models with a Girsanov-based KL regularizer, showing stability across different denoising step counts.

  15. CoDe: Blockwise Control for Denoising Diffusion Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CoDe applies blockwise best-of-N sampling during diffusion denoising, with Tweedie-based reward estimates, to align generated images to differentiable or non-differentiable rewards.

  16. Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

    cs.CV 2025-06 conditional novelty 4.0 of 10

    SmPO-Diffusion improves diffusion-model preference alignment with reward-model soft labels and ReNoise inversion, reporting higher human-preference scores and up to 26x lower training cost than Diffusion-KTO.

  17. Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

    cs.CV 2025-05 reject novelty 4.0 of 10

    Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.

  18. Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review

    cs.AI 2025-01 conditional novelty 3.0 of 10

    A tutorial showing that most inference-time reward-guided diffusion sampling methods approximate the same soft-optimal denoising policy, with some new algorithmic variants.

Pith tools