REVIEW 43 cited by
P+: Extended Textual Conditioning in Text-to-Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
We introduce an Extended Textual Conditioning space in text-to-image models, referred to as $P+$. This space consists of multiple textual conditions, derived from per-layer prompts, each corresponding to a layer of the denoising U-net of the diffusion model. We show that the extended space provides greater disentangling and control over image synthesis. We further introduce Extended Textual Inversion (XTI), where the images are inverted into $P+$, and represented by per-layer tokens. We show that XTI is more expressive and precise, and converges faster than the original Textual Inversion (TI) space. The extended inversion method does not involve any noticeable trade-off between reconstruction and editability and induces more regular inversions. We conduct a series of extensive experiments to analyze and understand the properties of the new space, and to showcase the effectiveness of our method for personalizing text-to-image models. Furthermore, we utilize the unique properties of this space to achieve previously unattainable results in object-style mixing using text-to-image models. Project page: https://prompt-plus.github.io
Forward citations
Cited by 43 Pith papers
-
LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers
LoRAShop localizes each LoRA's effect to attention-derived spatial masks inside a Flux transformer, enabling training-free multi-concept image generation and editing.
-
DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching
DreamCache achieves zero-shot personalized image generation by caching reference features from one denoising step and injecting them through 25M-parameter adapters.
-
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
Diptych Prompting generates images of a reference subject in new contexts by framing the task as text-conditioned inpainting of the right panel of a two-panel image, with no fine-tuning.
-
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
Adding persistently updated, supervised world-state register tokens to streaming multi-agent diffusion improves cross-agent consistency and visual quality in two-agent Minecraft generation.
-
PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation
Two-stage text-guided mesh deformation (Laplacian CLIP scaling + attention-shared SDS Jacobian sculpting) better preserves source pose while aligning to text than TextDeformer or MeshUp.
-
LILAC: Layer-Wise Independent LoRAs and Cascaded Conditioning for Multi-Concept Customization of Diffusion Models
Independently trained LoRAs composed as sequential layers with frozen conditioning preserve multi-subject identity better than weight-space fusion, reaching 0.861 ArcFace detection rate.
-
Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
Implicit generative choices in diffusion models for ambiguous prompts are localized principally in self-attention layers, enabling a targeted ICM steering method that outperforms prior debiasing approaches.
-
GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
GimmBO uses preference-based Bayesian optimization with a sparse, sum-bounded search space to help users interactively discover adapter merges in 20-30 dimensional model-merging spaces.
-
Alterbute: Editing Intrinsic Attributes of Objects in Images
Alterbute performs identity-preserving editing of an object's intrinsic attributes (color, texture, material, shape) using Visual-Named-Entity-based identity supervision and a relaxed training objective.
-
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
D-GPTScore, which averages GPT-4o's per-aspect ratings of concept-customized images, correlates with human preference at 0.78 Pearson on the new CC-AlignBench, beating prior metrics.
-
Per-Query Visual Concept Learning
A prompt- and seed-specific, attention-based loss step improves both identity preservation and prompt adherence for six personalization methods across SD, SDXL, and FLUX backbones.
-
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Layout-Togglable storytelling is introduced: diffusion transformers conditioned on layout enable precise control over character position and appearance, supported by a new large-scale dataset and benchmark.
-
TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
TARA adds token-focused masking and a token alignment loss to LoRA adapters, allowing several independently trained personalized adapters to be composed with less identity loss and feature leakage.
-
SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
SCFlow learns a reversible style-content merge and then lets the same mapping perform separation without explicit disentanglement training.
-
ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
ScenePainter introduces a SceneConceptGraph that encodes multi-level scene concepts and relations, and aligns an outpainting model with them to reduce semantic drift in perpetual 3D scene generation.
-
APT: Adaptive Personalized Training for Diffusion Models with Limited Data
APT detects overfitting during diffusion fine-tuning and uses adaptive augmentation, loss weighting, feature-statistics regularization, and attention alignment to preserve prior knowledge while learning new concepts.
-
Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
Difference Inversion learns text tokens encoding the A-to-A' edit and applies those tokens to B to synthesize B' with Stable Diffusion, without model-specific tuning.
-
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
Adding consistency-to-pretrained and multiplicative-noise consistency losses to fine-tuning improves subject identity and background diversity over DreamBooth on a 30-subject benchmark.
-
MARBLE: Material Recomposition and Blending in CLIP-Space
MARBLE performs material blending and parametric material-attribute control by manipulating CLIP image embeddings and injecting them into a specific U-Net block of a pre-trained diffusion model.
-
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
AlginGen improves zero-shot personalized image generation by training a learnable token and a selective attention mask that align textual and visual priors, achieving the best balance of concept preservation and promp...
-
CDST: Color Disentangled Style Transfer for Universal Style Reference Customization
CDST disentangles color from style via greyscale style input and a color histogram stream, enabling zero-shot style transfer with separate color control and a new characteristics-preserved mode.
-
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
VGD performs gradient-free hard prompt inversion by decoding with an LLM and reranking beams with CLIP similarity to the target image.
-
MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
MUSAR trains multi-subject text-to-image customization from a single-subject dataset by synthesizing diptych pairs and routing each image region's attention to the correct reference subject.
-
StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models
StyleBlend learns few-shot artistic style as separate layout and texture components and blends them during diffusion sampling to improve text-aligned, style-specific image generation.
-
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
Concatenating all frame prompts into a single prompt, then reweighting singular values and re-anchoring cross-attention, yields training-free identity-consistent text-to-image generation.
-
SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
SceneBooth keeps a provided subject image untouched and paints a new background around it, guided by a caption, object labels, and a predicted scene layout.
-
Nested Attention: Semantic-aware Attention Values for Concept Personalization
Nested Attention replaces a subject token's cross-attention value with a query-dependent value computed by an inner attention layer over image tokens, improving identity preservation and prompt adherence.
-
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
Omni-ID is a fixed-size, multi-view face representation trained with few-to-many reconstruction that reports higher identity preservation than ArcFace and CLIP in face generation and personalized text-to-image tasks.
-
LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
A hypernetwork pretrained on pairs of subject and style LoRAs predicts column-wise merging coefficients, enabling real-time, high-quality joint subject-style image personalization.
-
DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models
DreamBlend guides an overfit fine-tuned checkpoint with cross-attention maps from an underfit checkpoint, improving subject fidelity, prompt fidelity, and diversity in personalized text-to-image generation.
-
PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion
PersonaCraft adds SMPLx depth and normal conditioning, occlusion boundary enhancement, and occlusion-aware classifier-free guidance to diffusion models, enabling controllable multi-person images that preserve both fac...
-
HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation
HypDAE uses a hyperbolic latent space on top of a Stable Diffusion autoencoder so that few-shot image generation can vary identity-irrelevant details while a user-adjustable radius controls semantic diversity.
-
Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling
A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.
-
Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition
Pulsatile local energy deposition can initiate sustainable oblique detonation on a finite wedge with less than 10% of the average power needed by continuous deposition.
-
From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation
Wardrobe Polyptych LoRA lets a single diffusion model compose a person's face and clothing from multiple reference photos into new full-body images, generalizing to unseen identities without inference-time fine-tuning.
-
ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation
ShowFlow introduces KronA-WED adapter with SAR for single-concept and reuses it with SAMA plus layout guidance for condition-free multi-concept image generation.
-
Parallel Rescaling: Rebalancing Consistency Guidance for Personalized Diffusion Models
Re-centering and re-scaling the consistency guidance component parallel to the text direction improves prompt adherence with only a small drop in identity preservation.
-
Training Free Stylized Abstraction
A training-free framework coupling VLLM-based identity distillation with cross-domain rectified flow inversion generates identity-preserving stylized abstractions from a single reference image, evaluated by a new GPT-...
-
Preliminary Explorations with GPT-4o(mni) Native Image Generation
A qualitative exploration showing GPT-4o image generation excels at stylization, editing, and personalization but struggles with spatial reasoning, knowledge-based accuracy, and temporal prediction.
-
AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation
AnyStory introduces a unified feed-forward approach for single and multi-subject text-to-image personalization using a simplified ReferenceNet and CLIP encoder, plus a decoupled instance-aware router.
-
$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models
Diffusion model features, when decoded with k-sparse autoencoders, reveal interpretable visual concepts, and a lightweight classifier on the best layer (up_ft1 at t=25) beats prior diffusion-based classifiers on fine-...
-
Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attacks on Breast Ultrasound Images
Prompt2Perturb finds Stable Diffusion text embeddings that turn breast ultrasound images into adversarial examples that are natural-looking and mislead classifiers.
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.