Pith. sign in

REVIEW 30 cited by

P+: Extended Textual Conditioning in Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.09522 v3 pith:3L3SIAOR submitted 2023-03-16 cs.CV cs.CLcs.GRcs.LG

P+: Extended Textual Conditioning in Text-to-Image Generation

classification cs.CV cs.CLcs.GRcs.LG
keywords spaceextendedtextualtext-to-imageinversionmodelsconditioningintroduce
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce an Extended Textual Conditioning space in text-to-image models, referred to as $P+$. This space consists of multiple textual conditions, derived from per-layer prompts, each corresponding to a layer of the denoising U-net of the diffusion model. We show that the extended space provides greater disentangling and control over image synthesis. We further introduce Extended Textual Inversion (XTI), where the images are inverted into $P+$, and represented by per-layer tokens. We show that XTI is more expressive and precise, and converges faster than the original Textual Inversion (TI) space. The extended inversion method does not involve any noticeable trade-off between reconstruction and editability and induces more regular inversions. We conduct a series of extensive experiments to analyze and understand the properties of the new space, and to showcase the effectiveness of our method for personalizing text-to-image models. Furthermore, we utilize the unique properties of this space to achieve previously unattainable results in object-style mixing using text-to-image models. Project page: https://prompt-plus.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Inline Critic Steers Image Editing

    cs.CV 2026-05 conditional novelty 7.0

    Inline Critic uses a learnable token to critique and steer a frozen image-editing model's intermediate layers during generation, delivering state-of-the-art results on GEdit-Bench, RISEBench, and KRIS-Bench.

  2. PromptEvolver: Prompt Inversion through Evolutionary Optimization in Natural-Language Space

    cs.LG 2026-04 unverdicted novelty 7.0

    PromptEvolver recovers high-fidelity natural language prompts for given images by evolving them via genetic algorithm guided by a vision-language model, outperforming prior methods on benchmarks.

  3. DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

    cs.CV 2026-03 unverdicted novelty 7.0

    DSH-Bench is a benchmark for subject-driven T2I generation that uses hierarchical taxonomy sampling, difficulty/scenario classification, and a new SICS metric showing 9.4% higher human correlation than prior measures.

  4. DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

    cs.CV 2026-03 conditional novelty 6.5

    DSH-Bench supplies a hierarchical 58-category subject set, difficulty/scenario labels, and a human-aligned SICS metric that exposes systematic failures of 19 subject-driven T2I models.

  5. PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation

    cs.GR 2026-07 conditional novelty 6.0

    Two-stage text-guided mesh deformation (Laplacian CLIP scaling + attention-shared SDS Jacobian sculpting) better preserves source pose while aligning to text than TextDeformer or MeshUp.

  6. LILAC: Layer-Wise Independent LoRAs and Cascaded Conditioning for Multi-Concept Customization of Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0

    Independently trained LoRAs composed as sequential layers with frozen conditioning preserve multi-subject identity better than weight-space fusion, reaching 0.861 ArcFace detection rate.

  7. IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    IREU improves identity unlearning in CPG by offline location of identity features followed by targeted perturbations, outperforming global updates while preserving fidelity for retained identities and generalizing acr...

  8. Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting

    cs.CV 2026-06 unverdicted novelty 6.0

    Prompt-aware weighting strategies W-Switch and W-Composite improve multi-concept LoRA composition in diffusion models without training.

  9. Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

    cs.CV 2026-06 unverdicted novelty 6.0

    Equilibrated Diffusion decomposes concepts in frequency space to independently optimize subject and style embeddings, plus mask-guided diffusion and residual reference attention, for improved subject fidelity and text...

  10. SlimDiffSR: Toward Lightweight and Efficient Remote Sensing Image Super-Resolution via Diffusion Model Distillation

    cs.CV 2026-05 unverdicted novelty 6.0

    SlimDiffSR uses uncertainty-guided timestep assignment and structured pruning with frequency- and direction-separable convolutions plus MMD distillation to create a 200x faster, 20x smaller diffusion SR model for remo...

  11. PostureObjectstitch: Anomaly Image Generation Considering Assembly Relationships in Industrial Scenarios

    cs.CV 2026-04 unverdicted novelty 6.0

    PostureObjectStitch generates assembly-aware anomaly images by decoupling multi-view features into high-frequency, texture and RGB components, modulating them temporally in a diffusion model, and applying conditional ...

  12. Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models

    cs.CV 2026-03 unverdicted novelty 6.0

    Implicit generative choices in diffusion models for ambiguous prompts are localized principally in self-attention layers, enabling a targeted ICM steering method that outperforms prior debiasing approaches.

  13. NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion

    cs.CV 2025-11 unverdicted novelty 6.0

    NP-LoRA fuses subject and style LoRAs via null-space projection of the content update onto the orthogonal complement of the style subspace, with a soft variant controlled by one parameter.

  14. Adversarial Concept Distillation for One-Step Diffusion Personalization

    cs.CV 2025-10 unverdicted novelty 6.0

    OPAD enables reliable high-quality personalization of one-step diffusion models via multi-step teacher distillation combined with adversarial alignment losses.

  15. DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

    cs.SD 2025-09 unverdicted novelty 6.0

    DreamAudio generates audio clips that incorporate user-specified personalized audio events from reference samples while remaining aligned with text prompts.

  16. OmniPrism: Learning Disentangled Visual Concept for Image Generation

    cs.CV 2024-12 unverdicted novelty 6.0

    OmniPrism proposes a disentanglement method using a new paired dataset (PCD-200K), COD contrastive training, and block embeddings to inject separated concepts into diffusion models for multi-aspect image generation.

  17. PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation

    cs.GR 2026-07 conditional novelty 5.0

    PoseAlign splits text-guided mesh deformation into Laplacian-based global pose scaling and attention-sharing SDS local sculpting to keep pose while matching text.

  18. DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing

    cs.CV 2026-05 unverdicted novelty 5.0

    DreamEdit3D learns separate token embeddings for segmented object components via two-phase multi-view optimization to enable text-guided 3D editing with consistent image generation and mesh reconstruction.

  19. SlimDiffSR: Toward Lightweight and Efficient Remote Sensing Image Super-Resolution via Diffusion Model Distillation

    cs.CV 2026-05 unverdicted novelty 5.0

    SlimDiffSR uses uncertainty-guided timestep assignment and remote-sensing-tailored pruning plus MMD distillation to create a diffusion SR model with 200x faster inference and 20x fewer parameters while keeping competi...

  20. FREE-Switch: Frequency-based Dynamic LoRA Switch for Style Transfer

    cs.CV 2026-04 unverdicted novelty 5.0

    FREE-Switch dynamically switches LoRA adapters using frequency importance per diffusion step and adds semantic alignment to reduce content drift when merging specialized image generators.

  21. MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping

    cs.CV 2026-04 unverdicted novelty 5.0

    A scalable pipeline generates an intra-consistent, inter-diverse 1.4M style image dataset from text-to-image models and uses it to train a style encoder and generalizable style transfer model.

  22. Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models

    cs.CV 2026-03 unverdicted novelty 5.0

    Implicit generative choices in diffusion models concentrate in self-attention layers; targeted ICM interventions there outperform broader debiasing methods with fewer artifacts.

  23. PureCC: Pure Learning for Text-to-Image Concept Customization

    cs.CV 2026-03 unverdicted novelty 5.0

    PureCC introduces a decoupled learning objective, dual-branch training pipeline with frozen extractor, and adaptive guidance scale λ* for high-fidelity concept customization while preserving original model behavior in...

  24. TPGDiff: Hierarchical Triple-Prior Guided Diffusion for Image Restoration

    cs.CV 2026-01 unverdicted novelty 5.0

    TPGDiff introduces hierarchical triple-prior guidance in a diffusion network, placing degradation priors throughout, structural priors in shallow layers, and semantic priors in deep layers for improved all-in-one imag...

  25. SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

    cs.CV 2025-06 unverdicted novelty 5.0

    SynMotion combines disentangled semantic embeddings, parameter-efficient motion adapters, and alternate subject-motion training on a new SPV dataset to improve motion customization in text-to-video and image-to-video ...

  26. FA-Seg: A Fast and Accurate Diffusion-Based Method for Open-Vocabulary Segmentation

    cs.CV 2025-06 unverdicted novelty 5.0

    FA-Seg delivers state-of-the-art training-free open-vocabulary segmentation performance (43.8% mIoU average) on standard benchmarks by extracting and refining attention from a single forward pass of a pretrained diffu...

  27. ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation

    cs.CV 2025-06 unverdicted novelty 5.0

    ShowFlow introduces KronA-WED adapter with SAR for single-concept and reuses it with SAMA plus layout guidance for condition-free multi-concept image generation.

  28. Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift

    cs.CV 2025-05 unverdicted novelty 5.0

    Proposes Lipschitz regularization during fine-tuning to prevent distributional drift in personalized diffusion models, improving subject fidelity and prompt adherence.

  29. Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation

    cs.CV 2026-06 unverdicted novelty 4.0

    Early DC component convergence in text-to-image Transformer features causes output homogeneity; selective early attenuation via DAVE improves diversity without retraining or extra cost.

  30. TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation

    cs.CV 2024-09 unverdicted novelty 4.0

    TextBoost is a one-shot personalization technique that selectively fine-tunes the text encoder of diffusion models using causality-preserving adaptation and lightweight adapters to reduce parameters and storage.