REVIEW 5 cited by
Conditional Image Synthesis with Diffusion Models: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Conditional image synthesis based on user-specified requirements is a key component in creating complex visual content. In recent years, diffusion-based generative modeling has become a highly effective way for conditional image synthesis, leading to exponential growth in the literature. However, the complexity of diffusion-based modeling, the wide range of image synthesis tasks, and the diversity of conditioning mechanisms present significant challenges for researchers to keep up with rapid developments and to understand the core concepts on this topic. In this survey, we categorize existing works based on how conditions are integrated into the two fundamental components of diffusion-based modeling, $\textit{i.e.}$, the denoising network and the sampling process. We specifically highlight the underlying principles, advantages, and potential challenges of various conditioning approaches during the training, re-purposing, and specialization stages to construct a desired denoising network. We also summarize six mainstream conditioning mechanisms in the sampling process. All discussions are centered around popular applications. Finally, we pinpoint several critical yet still unsolved problems and suggest some possible solutions for future research. Our reviewed works are itemized at https://github.com/zju-pi/Awesome-Conditional-Diffusion-Models.
Forward citations
Cited by 5 Pith papers
-
Using Diffusion Models to do Data Assimilation
Diffusion DA systems with climatological, cycled, or forecast-augmented priors target different posterior distributions; only a per-cycle retrained model matches ensemble DA.
-
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Appearance pointers are compact tokens that let a diffusion transformer apply text, image, or combined prompts to specific image regions in a single pass.
-
Generative Latent Diffusion for Efficient Spatiotemporal Data Reduction
A latent diffusion model conditioned on keyframe latents reconstructs non-key frames, giving higher compression ratios than prior scientific data compressors.
-
Translationese as a Rational Response to Translation Task Difficulty
Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.
-
Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance
Adding InfoNCE loss and a SegFormer mask encoder to a Giambi-style face diffusion pipeline lowers the reported FID from 74.07 to 70.98, and to 63.85 when segmentation is included.
Discussion (0). Sign in to comment.