Pith. sign in

REVIEW 7 cited by

Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.02642 v1 pith:EBBIXWKD submitted 2023-04-05 cs.CV

classification cs.CV
keywords embeddinggenerationobjectmethodobjectsoptimizationtext-to-imageable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes a method for generating images of customized objects specified by users. The method is based on a general framework that bypasses the lengthy optimization required by previous approaches, which often employ a per-object optimization paradigm. Our framework adopts an encoder to capture high-level identifiable semantics of objects, producing an object-specific embedding with only a single feed-forward pass. The acquired object embedding is then passed to a text-to-image synthesis model for subsequent generation. To effectively blend a object-aware embedding space into a well developed text-to-image model under the same generation context, we investigate different network designs and training strategies, and propose a simple yet effective regularized joint training scheme with an object identity preservation loss. Additionally, we propose a caption generation scheme that become a critical piece in fostering object specific embedding faithfully reflected into the generation process, while keeping control and editing abilities. Once trained, the network is able to produce diverse content and styles, conditioned on both texts and objects. We demonstrate through experiments that our proposed method is able to synthesize images with compelling output quality, appearance diversity, and object fidelity, without the need of test-time optimization. Systematic studies are also conducted to analyze our models, providing insights for future work.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

    cs.CV 2023-07 unverdicted novelty 7.0 of 10

    A single motion module trained on videos adds temporally coherent animation to any personalized text-to-image model derived from the same base without additional tuning.

  2. Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Equilibrated Diffusion decomposes concepts in frequency space to independently optimize subject and style embeddings, plus mask-guided diffusion and residual reference attention, for improved subject fidelity and text...

  3. Intrinsic Concept Extraction Based on Compositional Interpretability

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    HyperExpress extracts composable intrinsic concepts from single images via hyperbolic concept learning and concept-wise optimization in diffusion-based models.

  4. Adversarial Concept Distillation for One-Step Diffusion Personalization

    cs.CV 2025-10 unverdicted novelty 6.0 of 10

    OPAD enables reliable high-quality personalization of one-step diffusion models via multi-step teacher distillation combined with adversarial alignment losses.

  5. Per-Query Visual Concept Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A prompt- and seed-specific, attention-based loss step improves both identity preservation and prompt adherence for six personalization methods across SD, SDXL, and FLUX backbones.

  6. Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Contrastive Inversion disentangles the common concept from per-image auxiliary tokens via InfoNCE loss, then fine-tunes only the target cross-attention pathway, matching DisenBooth's numbers while claiming better qual...

  7. Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A multi-view diffusion hair-transfer system transfers a reference hairstyle onto a portrait and renders the edited person from many consistent viewpoints.

Pith tools