A lightweight transformer module learns quality-aware vectors from timestep and prompt embeddings to modulate adaptive LayerNorm in DiT blocks, yielding consistent image quality gains over baseline diffusion transformers.
Reward- instruct: A reward-centric approach to fast photo-realistic image generation.arXiv preprint arXiv:2503.13070
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3verdicts
UNVERDICTED 3roles
background 1polarities
background 1representative citing papers
Sol-RL decouples FP4-based candidate exploration from BF16 policy optimization in diffusion RL, delivering up to 4.64x faster convergence with maintained or superior alignment performance on models like FLUX.1 and SD3.5.
Reward-Forcing guides autoregressive video generation with reward feedback to achieve performance comparable to teacher-dependent methods on benchmarks like VBench without relying on distillation.
citing papers explorer
-
Quality-Aware Modulation for Diffusion Transformers
A lightweight transformer module learns quality-aware vectors from timestep and prompt embeddings to modulate adaptive LayerNorm in DiT blocks, yielding consistent image quality gains over baseline diffusion transformers.
-
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
Sol-RL decouples FP4-based candidate exploration from BF16 policy optimization in diffusion RL, delivering up to 4.64x faster convergence with maintained or superior alignment performance on models like FLUX.1 and SD3.5.
-
Reward-Forcing: Autoregressive Video Generation with Reward Feedback
Reward-Forcing guides autoregressive video generation with reward feedback to achieve performance comparable to teacher-dependent methods on benchmarks like VBench without relying on distillation.