Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline

· 2026 · cs.AI · arXiv 2605.06583

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

We propose a deterministic adjoint matching framework that formulates human preference alignment for flow-based generative models as an optimal control problem over velocity fields. One can directly regress the control toward a value-gradient-induced target under the current policy, leading to a simple and stable training objective. Building on this perspective, we introduce a truncated adjoint scheme that focuses computation on the terminal portion of the trajectory, where reward-relevant signals concentrate, which yields substantial computational savings while preserving alignment quality. We further generalize the framework beyond standard KL-based regularization, allowing more flexible trade-offs between alignment strength and distributional preservation. Experiments on SiT-XL/2 and FLUX.2-Klein-4B demonstrate consistent gains across multiple alignment metrics, along with substantially improved diversity and mode preservation.

representative citing papers

OPD+: Rethinking the Advantage Design for On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 7.0

OPD+ removes the bias from stop-gradient in on-policy distillation by deriving correct gradients for f-divergences, outperforming standard KL-based methods on math reasoning and tool-use tasks.

citing papers explorer

Showing 1 of 1 citing paper after filters.

OPD+: Rethinking the Advantage Design for On-Policy Distillation cs.LG · 2026-05-31 · unverdicted · none · ref 4 · internal anchor
OPD+ removes the bias from stop-gradient in on-policy distillation by deriving correct gradients for f-divergences, outperforming standard KL-based methods on math reasoning and tool-use tasks.

Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline

fields

years

verdicts

representative citing papers

citing papers explorer