Pith. sign in

REVIEW 3 cited by

PixelHacker: Image Inpainting with Structural and Semantic Consistency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.20438 v2 pith:EQSNC6JY submitted 2025-04-29 cs.CV

classification cs.CV
keywords pixelhackerimageconsistencyinpaintingattentionbackgroundcategoriesdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image inpainting is a fundamental research area between image editing and image generation. Recent state-of-the-art (SOTA) methods have explored novel attention mechanisms, lightweight architectures, and context-aware modeling, demonstrating impressive performance. However, they often struggle with complex structure (e.g., texture, shape, spatial relations) and semantics (e.g., color consistency, object restoration, and logical correctness), leading to artifacts and inappropriate generation. To address this challenge, we design a simple yet effective inpainting paradigm called latent categories guidance, and further propose a diffusion-based model named PixelHacker. Specifically, we first construct a large dataset containing 14 million image-mask pairs by annotating foreground and background (potential 116 and 21 categories, respectively). Then, we encode potential foreground and background representations separately through two fixed-size embeddings, and intermittently inject these features into the denoising process via linear attention. Finally, by pre-training on our dataset and fine-tuning on open-source benchmarks, we obtain PixelHacker. Extensive experiments show that PixelHacker comprehensively outperforms the SOTA on a wide range of datasets (Places2, CelebA-HQ, and FFHQ) and exhibits remarkable consistency in both structure and semantics. Project page at https://hustvl.github.io/PixelHacker.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Moebius introduces a compressed diffusion inpainting model using Local-λ Mix Interaction blocks and latent-space multi-granularity distillation to reach 10B-level quality with 0.22B parameters.

  2. ClickRemoval: An Interactive Open-Source Tool for Object Removal in Diffusion Models

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ClickRemoval delivers click-driven object removal and background restoration in diffusion models through self-attention modulation without additional training or inputs.

  3. 2D Instance Editing in 3D Space

    cs.CV 2025-07 reject novelty 4.0 of 10

    A 2D-to-3D-to-2D editing system that segments an object, reconstructs it as 3D Gaussians, deforms it under a rigidity constraint, and inpaints it back into the original image.

Pith tools