Pith. sign in

REVIEW 5 cited by

DreamInpainter: Text-Guided Subject-Driven Image Inpainting with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03771 v1 pith:TEJX3GSW submitted 2023-12-05 cs.CV

classification cs.CV
keywords subjecttextimageinpaintingexemplarimagessubject-driventext-guided
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study introduces Text-Guided Subject-Driven Image Inpainting, a novel task that combines text and exemplar images for image inpainting. While both text and exemplar images have been used independently in previous efforts, their combined utilization remains unexplored. Simultaneously accommodating both conditions poses a significant challenge due to the inherent balance required between editability and subject fidelity. To tackle this challenge, we propose a two-step approach DreamInpainter. First, we compute dense subject features to ensure accurate subject replication. Then, we employ a discriminative token selection module to eliminate redundant subject details, preserving the subject's identity while allowing changes according to other conditions such as mask shape and text prompts. Additionally, we introduce a decoupling regularization technique to enhance text control in the presence of exemplar images. Our extensive experiments demonstrate the superior performance of our method in terms of visual quality, identity preservation, and text control, showcasing its effectiveness in the context of text-guided subject-driven image inpainting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BrushEdit: All-In-One Image Inpainting and Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.

  2. PainterNet: Adaptive Image Inpainting with Actual-Token Attention and Diverse Mask Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PainterNet is a diffusion-model plugin that uses local prompts, attention supervision, and diverse masks to improve text-consistent image inpainting.

  3. InsightEdit: Towards Better Instruction Following for Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    InsightEdit uses multimodal language model features in a two-stream adapter to improve complex instruction following and background consistency in image editing.

  4. Generating Compositional Scenes via Text-to-image RGBA Instance Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A multi-stage text-to-image approach that generates individual objects as RGBA images and composes them scene-by-scene via noise blending, enabling fine-grained layout and attribute control.

  5. SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts

    cs.CR 2025-07 reject novelty 4.0 of 10

    A diffusion editing model is fine-tuned with a blur target for forbidden images and the original output for permitted images, claiming selective suppression of unauthorized edits.

Pith tools