Pith. sign in

REVIEW 19 cited by

A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03594 v4 pith:GQG7S75V submitted 2023-12-06 cs.CV

classification cs.CV
keywords inpaintingpowerpainttaskobjectfillingmodelpromptprompts
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Advancing image inpainting is challenging as it requires filling user-specified regions for various intents, such as background filling and object synthesis. Existing approaches focus on either context-aware filling or object synthesis using text descriptions. However, achieving both tasks simultaneously is challenging due to differing training strategies. To overcome this challenge, we introduce PowerPaint, the first high-quality and versatile inpainting model that excels in multiple inpainting tasks. First, we introduce learnable task prompts along with tailored fine-tuning strategies to guide the model's focus on different inpainting targets explicitly. This enables PowerPaint to accomplish various inpainting tasks by utilizing different task prompts, resulting in state-of-the-art performance. Second, we demonstrate the versatility of the task prompt in PowerPaint by showcasing its effectiveness as a negative prompt for object removal. Moreover, we leverage prompt interpolation techniques to enable controllable shape-guided object inpainting, enhancing the model's applicability in shape-guided applications. Finally, we conduct extensive experiments and applications to verify the effectiveness of PowerPaint. We release our codes and models on our project page: https://powerpaint.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.

  2. Score-based diffusion models for accurate crystal-structure inpainting and reconstruction of hydrogen positions

    cond-mat.mtrl-sci 2026-01 conditional novelty 6.0 of 10

    Adapting TD-Paint to crystal diffusion models reconstructs hydrogen positions with a LES success rate above 97%, beating unconditioned diffusion and DFT-based inpainting.

  3. Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-Planner, an MLLM-based planner, decomposes complex image-editing instructions into localized sub-edits with masks and boxes, improving editing quality on standard and new complex benchmarks.

  4. DreamLight: Towards Harmonious and Consistent Image Relighting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A unified image- and text-based relighting model with direction-biased attention and a wavelet foreground fixer outperforms existing methods on a synthetic relighting benchmark.

  5. Towards Seamless Borders: A Method for Mitigating Inconsistencies in Image Inpainting and Outpainting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-step training loss plus a fine-tuned VAE reduces color and structure discontinuities at mask boundaries in diffusion image inpainting and outpainting.

  6. ColorFlow: Retrieval-Augmented Image Sequence Colorization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ColorFlow is a three-stage diffusion framework that colorizes black-and-white image sequences while preserving character and object color identity via retrieved reference patches.

  7. BrushEdit: All-In-One Image Inpainting and Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.

  8. Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A mask-and-inpaint pipeline generates localized gender counterfactuals that preserve image context and, after fine-tuning, reduce gender bias in CLIP models.

  9. Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned Inpainting

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Pinco is a plug-and-play adapter that enables diffusion transformers to inpaint backgrounds around a provided foreground object, preserving its shape via self-attention injection and a positional anchor.

  10. FreeCond: Free Lunch in the Input Conditions of Text-Guided Inpainting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FreeCond adjusts only the image and mask inputs of Stable Diffusion Inpainting, improving prompt adherence and mask fitting without training or extra compute.

  11. Generative Image Layer Decomposition with Visual Effects

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A diffusion-based model decomposes an image into a clean background and a transparent foreground layer that retains shadows and reflections, enabling object removal and spatial edits.

  12. InsightEdit: Towards Better Instruction Following for Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    InsightEdit uses multimodal language model features in a two-stream adapter to improve complex instruction following and background consistency in image editing.

  13. VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A large open-source hybrid image-video dataset and a LoRA-based diffusion baseline for interactive local video editing.

  14. DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D Inpainting

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DiGA3D performs text-guided 3D inpainting (removal, re-texturing, replacement) with a coarse-to-fine diffusion propagation scheme to improve multi-view appearance and geometry consistency.

  15. Instructive3D: Editing Large Reconstruction Models with Text Instructions

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A text-conditioned diffusion adapter operating on the triplane latents of a frozen large reconstruction model enables natural-language editing of generated 3D objects.

  16. RORem: Training a Robust Object Remover with Human-in-the-Loop

    cs.CV 2025-01 conditional novelty 5.0 of 10

    RORem trains an SDXL-based object remover on a 200K-pair dataset grown by iterative human feedback and a learned discriminator, surpassing prior methods by roughly 18 points in human-judged success rate.

  17. Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    FloED uses optical flow as extra motion guidance in a diffusion video inpainting model, with flow-warped latent interpolation and attention caching to cut inference cost while improving temporal consistency.

  18. MagicQuill: An Intelligent Interactive Image Editing System

    cs.CV 2024-11 conditional novelty 5.0 of 10

    MagicQuill combines brush-based edge and color control with an MLLM that guesses user intent, enabling fast interactive image edits without typing prompts.

  19. MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting

    cs.CV 2025-06 conditional novelty 4.0 of 10

    MTADiffusion improves text-guided object inpainting by training on a new 5M-image mask-text dataset with edge prediction and style-consistency losses.

Pith tools