D2PO learns better low-NFE diffusion timestep schedules and CFG weights via DPO on a score-based energy with a dynamic denser-schedule preference target.
hub
In: CVPR (2022)
14 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
StableMTL repurposes latent diffusion models for multi-task learning from partially annotated synthetic data via unified latent loss, task encoding, and a multi-stream task-attention architecture, reporting outperformance on 7 tasks across 8 benchmarks.
Pretrained VAE latents have far lower transmission–reflection coherence than pixels, so flow matching on FLUX plus cycle and contrastive losses can separate both layers in one pass.
FlowCIR frames ZS-CIR as conditional flow matching transport on fixed VLM embeddings plus an inference-time Multi-Negative Steering fix for negation, reporting competitive benchmark results at far lower training cost.
PanoGaussian distills panoramic representations into explicit dynamic Gaussians for consistent monocular 4D scene synthesis under large viewpoint variations.
Anti-Prompt protects images from text-guided image-to-video generation by suppressing text-conditioned attention during denoising, producing visible generation failures.
DriveWeaver performs point-conditioned video inpainting with a global-to-local hierarchical strategy to insert controllable vehicles into autonomous driving simulations and extracts 3D Gaussians for real-time rendering.
A saliency-guided warp-unwarp method reallocates spatial representation to preserve fine structures in latent diffusion models for image-to-image translation.
Mural transfers knowledge from a frozen LLM to text-to-image synthesis via MoT shared attention, achieving 0.85 GenEval, 86.75 DPG-Bench, and 0.66 WISE while exhibiting emergent behaviors without multimodal or reasoning supervision.
Refinement via Regeneration (RvR) reformulates image refinement in unified multimodal models as conditional regeneration using prompt and semantic tokens from the initial image, yielding higher alignment scores than editing-based methods.
HO-Flow synthesizes realistic hand-object motions from text and canonical 3D objects via an interaction-aware VAE and masked flow matching, reporting SOTA physical plausibility and diversity on GRAB, OakInk, and DexYCB.
MIRAGE introduces a benchmark for multi-instance image editing and a training-free framework that uses vision-language parsing and parallel regional denoising to achieve precise edits without altering backgrounds.
SpatialEdit provides a benchmark, large synthetic dataset, and baseline model for precise object and camera spatial manipulations in images, with the model beating priors on spatial editing.
ALM integrates likelihood maximization and acceleration into diffusion reverse sampling to enable globally coherent generation from incomplete inputs.
citing papers explorer
-
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
D2PO learns better low-NFE diffusion timestep schedules and CFG weights via DPO on a score-based energy with a dynamic denser-schedule preference target.
-
StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets
StableMTL repurposes latent diffusion models for multi-task learning from partially annotated synthetic data via unified latent loss, task encoding, and a multi-stream task-attention architecture, reporting outperformance on 7 tasks across 8 benchmarks.
-
PRISM: Latent Composition Consistency for Single-Image Reflection Removal
Pretrained VAE latents have far lower transmission–reflection coherence than pixels, so flow matching on FLUX plus cycle and contrastive losses can separate both layers in one pass.
-
FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
FlowCIR frames ZS-CIR as conditional flow matching transport on fixed VLM embeddings plus an inference-time Multi-Negative Steering fix for negation, reporting competitive benchmark results at far lower training cost.
-
Unified Panoramic-Gaussian Representation for Monocular 4D Scene Synthesis
PanoGaussian distills panoramic representations into explicit dynamic Gaussians for consistent monocular 4D scene synthesis under large viewpoint variations.
-
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
Anti-Prompt protects images from text-guided image-to-video generation by suppressing text-conditioned attention during denoising, producing visible generation failures.
-
DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation
DriveWeaver performs point-conditioned video inpainting with a global-to-local hierarchical strategy to insert controllable vehicles into autonomous driving simulations and extracts 3D Gaussians for real-time rendering.
-
WarpI2I: Image Warping for Image-to-Image Translation
A saliency-guided warp-unwarp method reallocates spatial representation to preserve fine structures in latent diffusion models for image-to-image translation.
-
Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers
Mural transfers knowledge from a frozen LLM to text-to-image synthesis via MoT shared attention, achieving 0.85 GenEval, 86.75 DPG-Bench, and 0.66 WISE while exhibiting emergent behaviors without multimodal or reasoning supervision.
-
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
Refinement via Regeneration (RvR) reformulates image refinement in unified multimodal models as conditional regeneration using prompt and semantic tokens from the initial image, yielding higher alignment scores than editing-based methods.
-
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
HO-Flow synthesizes realistic hand-object motions from text and canonical 3D objects via an interaction-aware VAE and masked flow matching, reporting SOTA physical plausibility and diversity on GRAB, OakInk, and DexYCB.
-
MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing
MIRAGE introduces a benchmark for multi-instance image editing and a training-free framework that uses vision-language parsing and parallel regional denoising to achieve precise edits without altering backgrounds.
-
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
SpatialEdit provides a benchmark, large synthetic dataset, and baseline model for precise object and camera spatial manipulations in images, with the model beating priors on spatial editing.
-
Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation
ALM integrates likelihood maximization and acceleration into diffusion reverse sampling to enable globally coherent generation from incomplete inputs.