REVIEW 5 cited by
SEGA: Instructing Text-to-Image Models using Semantic Guidance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Text-to-image diffusion models have recently received a lot of interest for their astonishing ability to produce high-fidelity images from text only. However, achieving one-shot generation that aligns with the user's intent is nearly impossible, yet small changes to the input prompt often result in very different images. This leaves the user with little semantic control. To put the user in control, we show how to interact with the diffusion process to flexibly steer it along semantic directions. This semantic guidance (SEGA) generalizes to any generative architecture using classifier-free guidance. More importantly, it allows for subtle and extensive edits, changes in composition and style, as well as optimizing the overall artistic conception. We demonstrate SEGA's effectiveness on both latent and pixel-based diffusion models such as Stable Diffusion, Paella, and DeepFloyd-IF using a variety of tasks, thus providing strong evidence for its versatility, flexibility, and improvements over existing methods.
Forward citations
Cited by 5 Pith papers
-
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
A zero-shot method combines grid priors, depth-conditioned ControlNet, and latent blending to generate and consistently edit story visualizations across multiple frames.
-
BrushEdit: All-In-One Image Inpainting and Editing
BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.
-
FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers
FluxSpace performs training-free, disentangled semantic editing in rectified flow transformers by combining attention outputs with prompt-derived linear directions.
-
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
Embedding-space perturbation sensitivity, combined with diversity and leave-one-out fairness scores, can flag and localize intentionally biased text-to-image models.
-
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
GTF is a training-free, projection-based noise composition rule that enables text-driven addition, removal, and style transfer in diffusion models across image, video, and 3D generation.
Discussion (0). Continue with ORCID to comment.