Pith. sign in

REVIEW 3 cited by

SEGA: Instructing Text-to-Image Models using Semantic Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12247 v2 pith:Y37DTZLL submitted 2023-01-28 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords diffusionsemanticguidancemodelssegauserchangescontrol
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image diffusion models have recently received a lot of interest for their astonishing ability to produce high-fidelity images from text only. However, achieving one-shot generation that aligns with the user's intent is nearly impossible, yet small changes to the input prompt often result in very different images. This leaves the user with little semantic control. To put the user in control, we show how to interact with the diffusion process to flexibly steer it along semantic directions. This semantic guidance (SEGA) generalizes to any generative architecture using classifier-free guidance. More importantly, it allows for subtle and extensive edits, changes in composition and style, as well as optimizing the overall artistic conception. We demonstrate SEGA's effectiveness on both latent and pixel-based diffusion models such as Stable Diffusion, Paella, and DeepFloyd-IF using a variety of tasks, thus providing strong evidence for its versatility, flexibility, and improvements over existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BrushEdit: All-In-One Image Inpainting and Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.

  2. FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FluxSpace performs training-free, disentangled semantic editing in rectified flow transformers by combining attention outputs with prompt-derived linear directions.

  3. On the Fairness, Diversity and Reliability of Text-to-Image Generative Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Embedding-space perturbation sensitivity, combined with diversity and leave-one-out fairness scores, can flag and localize intentionally biased text-to-image models.

Pith tools