STEDiff improves semantic alignment in text-to-image diffusion models via training-free embedding strengthening with the [EOT] token and a spatial semantic loss, showing gains on T2I-CompBench.
Uncovering the Text Embedding in Text-to-Image Diffusion Models
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
Token-to-Token alignment rephrases prompts into shared structure then matches token embeddings by semantic similarity, making linear interpolation a meaningful operation for blending in text-to-image models.
sep-CMA-ES outperforms Adam on a combined aesthetic-plus-alignment objective when optimizing prompt embeddings for Stable Diffusion XL Turbo across 36 Parti Prompts and three weight settings.
citing papers explorer
-
STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model
STEDiff improves semantic alignment in text-to-image diffusion models via training-free embedding strengthening with the [EOT] token and a spatial semantic loss, showing gains on T2I-CompBench.
-
Token-to-Token Alignment of Text Embeddings for Semantic Blending
Token-to-Token alignment rephrases prompts into shared structure then matches token embeddings by semantic similarity, making linear interpolation a meaningful operation for blending in text-to-image models.
-
Evolutionary Optimization Trumps Adam Optimization on Embedding Space Exploration
sep-CMA-ES outperforms Adam on a combined aesthetic-plus-alignment objective when optimizing prompt embeddings for Stable Diffusion XL Turbo across 36 Parti Prompts and three weight settings.