pith. sign in

hub Canonical reference

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Canonical reference. 93% of citing Pith papers cite this work as background.

98 Pith papers citing it
Background 93% of classified citations
abstract

Text-to-image models offer unprecedented freedom to guide creation through natural language. Yet, it is unclear how such freedom can be exercised to generate images of specific unique concepts, modify their appearance, or compose them in new roles and novel scenes. In other words, we ask: how can we use language-guided models to turn our cat into a painting, or imagine a new product based on our favorite toy? Here we present a simple approach that allows such creative freedom. Using only 3-5 images of a user-provided concept, like an object or a style, we learn to represent it through new "words" in the embedding space of a frozen text-to-image model. These "words" can be composed into natural language sentences, guiding personalized creation in an intuitive way. Notably, we find evidence that a single word embedding is sufficient for capturing unique and varied concepts. We compare our approach to a wide range of baselines, and demonstrate that it can more faithfully portray the concepts across a range of applications and tasks. Our code, data and new words will be available at: https://textual-inversion.github.io

hub tools

citation-role summary

background 15

citation-polarity summary

claims ledger

  • abstract Text-to-image models offer unprecedented freedom to guide creation through natural language. Yet, it is unclear how such freedom can be exercised to generate images of specific unique concepts, modify their appearance, or compose them in new roles and novel scenes. In other words, we ask: how can we use language-guided models to turn our cat into a painting, or imagine a new product based on our favorite toy? Here we present a simple approach that allows such creative freedom. Using only 3-5 images of a user-provided concept, like an object or a style, we learn to represent it through new "wor
  • background These Preprint. arXiv:2605.07257v1 [cs.CV] 8 May 2026 advances have fueled growing interest in generative personalization: adapting a pretrained T2I model to a user-specific concept (e.g., a person, pet, or object) from only a few reference images, while retaining the ability to place that concept into novel contexts via natural-language prompts [10, 30]. The core objective is to preserve the unique identity of the personal concept while remaining faithful to the prompt's semantics. Despite rapi
  • background tasks including image synthesis [4, 24, 28], 3D object gen- eration [16, 21], and video production [1, 11, 29]. Leverag- ing large-scale pre-training on massive datasets, these mod- els now outperform earlier approaches in producing high- fidelity and coherent generative content. Current approaches range from slow fine-tuning methods like DreamBooth [26] and Textual Inversion [6], to zero- shot ID injection with encoders like IP-Adapter [38], Pho- toMaker [15], and InstantID [36], but these sacr
  • background they frequently incur information loss in either foreground objects or background contexts. 2.2 Testing-Time Finetuning Testing-time finetuning methods constitute a fundamental para- digm for personalized image generation, where pre-trained model parameters are adaptively optimized for specific target subjects dur- ing inference to achieve high-fidelity customized image synthesis. Textual Inversion [11] first introduced the concept of optimizing the embeddings of learnable tokens by incorporatin
  • background Finally, the model is highly sensitive to the prompt, and small changes in wording can lead to drastically different generated images, while semantically equivalent prompts may yield very different visual outputs [9,29]. Recently, a substantial body of works has tackled prompt inversion through optimization in continuous embedding or latent spaces [11,30,36,46]. While these methods can achieve high-fidelity reconstruction, they suffer from several fun- damental limitations. First, they assume wh
  • background in the scene, while camera motion adjusts the camera's position and angle. 5.2.1 Motion Customization. Motion customization generates videos with motions matching reference videos, requiring disentanglement of motion and appearance. Customize-A-Video [210] utilizes Temporal LoRA (T-LoRA) to learn motion from temporal layers and Appearance Absorbers(e.g., spatial LoRA or textual inversion [211]) to isolate spatial features. MotionDi- rector [212] employs dual-path LoRA: spatial LoRAs capture appe
  • background articulate the desired target through text descriptions. For instance, it is difficult to describe the precise features of an innovative toy car which is not encountered during large-scale model training. Consequently, the objective of customized generation is to enable the model to grasp new concepts from a minimal set of user-supplied images. Textual Inversion [243] addresses this by finding a new pseudo-word S˚ (similar to soft prompt discussed in Section III-A2) that represents new, specific

co-cited works

clear filters

representative citing papers

ZIPP:Zero-shot Image Personalization from Personas

cs.AI · 2026-06-07 · unverdicted · novelty 7.0

ZIPP conditions diffusion models on LLM-rewritten prompts derived from graph-mined natural-language personas to achieve zero-shot personalization, reporting 13-20% gains and 79% human preference win rate over generic outputs.

Functionalization via Structure Completion and Motion Rectification

cs.CV · 2026-05-18 · unverdicted · novelty 7.0

Object functionalization is cast as neural graph completion over a functional graph of parts, contacts, and motions, followed by geometry realization that also rectifies erroneous motions, demonstrated on furniture with a new paired dataset.

Adaptive Subspace Projection for Generative Personalization

cs.CV · 2026-05-08 · unverdicted · novelty 7.0

A training-free adaptive subspace projection method mitigates semantic collapsing in generative personalization by isolating and adjusting drift in a low-dimensional subspace using the stable pre-trained embedding as anchor.

Image-Guided Geometric Stylization of 3D Meshes

cs.CV · 2026-04-09 · unverdicted · novelty 7.0

A coarse-to-fine pipeline deforms 3D meshes to reflect geometric features from an image using diffusion model representations while preserving topology and part-level semantics.

Learning Interactive Real-World Simulators

cs.AI · 2023-10-09 · conditional · novelty 7.0

UniSim learns a universal real-world simulator from orchestrated diverse datasets, enabling zero-shot deployment of policies trained purely in simulation.

citing papers explorer

Showing 12 of 12 citing papers after filters.

  • Adaptive Subspace Projection for Generative Personalization cs.CV · 2026-05-08 · unverdicted · none · ref 10 · internal anchor

    A training-free adaptive subspace projection method mitigates semantic collapsing in generative personalization by isolating and adjusting drift in a low-dimensional subspace using the stable pre-trained embedding as anchor.

  • Personalizing Text-to-Image Generation to Individual Taste cs.CV · 2026-04-08 · unverdicted · none · ref 16 · internal anchor

    PAMELA provides a multi-user rating dataset and personalized reward model that predicts individual image preferences more accurately than prior population-level aesthetic models.

  • AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning cs.CV · 2023-07-10 · unverdicted · none · ref 6 · internal anchor

    A single motion module trained on videos adds temporally coherent animation to any personalized text-to-image model derived from the same base without additional tuning.

  • D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models cs.CV · 2026-05-06 · unverdicted · none · ref 20 · 3 links · internal anchor

    D-OPSD formulates supervised fine-tuning of step-distilled diffusion models as on-policy self-distillation by having the model act as both teacher (with multimodal context) and student (with text-only context) on its own roll-outs.

  • PostureObjectstitch: Anomaly Image Generation Considering Assembly Relationships in Industrial Scenarios cs.CV · 2026-04-15 · unverdicted · none · ref 11 · internal anchor

    PostureObjectStitch generates assembly-aware anomaly images by decoupling multi-view features into high-frequency, texture and RGB components, modulating them temporally in a diffusion model, and applying conditional loss plus geometric priors to preserve correct component relationships.

  • SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing cs.CV · 2026-04-06 · unverdicted · none · ref 14 · internal anchor

    SpatialEdit provides a benchmark, large synthetic dataset, and baseline model for precise object and camera spatial manipulations in images, with the model beating priors on spatial editing.

  • IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models cs.CV · 2023-08-13 · unverdicted · none · ref 51 · internal anchor

    IP-Adapter adds effective image prompting to text-to-image diffusion models using a lightweight decoupled cross-attention adapter that works alongside text prompts and other controls.

  • eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers cs.CV · 2022-11-02 · unverdicted · none · ref 18 · internal anchor

    An ensemble of stage-specialized text-to-image diffusion models improves prompt alignment over single shared-parameter models while preserving visual quality and inference speed.

  • RealDiffusion: Physics-informed Attention for Multi-character Storybook Generation cs.CV · 2026-05-12 · unverdicted · none · ref 6 · internal anchor

    RealDiffusion uses heat diffusion as a dissipative prior and a region-aware stochastic process inside a training-free physics-informed attention mechanism to improve multi-character coherence while preserving narrative dynamism in sequential image generation.

  • ID-Sim: An Identity-Focused Similarity Metric cs.CV · 2026-04-06 · unverdicted · none · ref 19 · internal anchor

    ID-Sim is a new similarity metric that aims to capture human selective sensitivity to identities by training on curated real and generative synthetic data and validating against human annotations on recognition, retrieval, and generative tasks.

  • LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cs.CV · 2026-04-13 · unverdicted · none · ref 45 · internal anchor

    This review organizes literature on large multimodal models and object-centric vision into four themes—understanding, referring segmentation, editing, and generation—while summarizing paradigms, strategies, and challenges like instance permanence and consistent interaction.

  • Evolution of Video Generative Foundations cs.CV · 2026-04-07 · unverdicted · none · ref 211 · internal anchor

    This survey traces video generation technology from GANs to diffusion models and then to autoregressive and multimodal approaches while analyzing principles, strengths, and future trends.