Pith. sign in

REVIEW 3 cited by

Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.02024 v3 pith:SEYOM6AW submitted 2022-12-05 cs.CV cs.LG

classification cs.CVcs.LG
keywords editingguidanceimageeditedpixel-wiserequirementsdiffusionfine-grained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our goal is to develop fine-grained real-image editing methods suitable for real-world applications. In this paper, we first summarize four requirements for these methods and propose a novel diffusion-based image editing framework with pixel-wise guidance that satisfies these requirements. Specifically, we train pixel-classifiers with a few annotated data and then infer the segmentation map of a target image. Users then manipulate the map to instruct how the image will be edited. We utilize a pre-trained diffusion model to generate edited images aligned with the user's intention with pixel-wise guidance. The effective combination of proposed guidance and other techniques enables highly controllable editing with preserving the outside of the edited area, which results in meeting our requirements. The experimental results demonstrate that our proposal outperforms the GAN-based method for editing quality and speed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage diffusion model generates realistic depth noise on synthetic CAD data, and pretraining 3D networks on the resulting data improves few-shot real-world 3D tasks.

  2. Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

    cs.CV 2025-06 reject novelty 5.0 of 10

    A new micro-edit dataset and fine-tuning recipe appear to help multimodal LLMs notice small visual changes, but the central 'feature consistency loss' claim is not present in the method.

  3. TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A diffusion model trained progressively from Kingdom to Species generates more accurate fine-grained animal images, including rare species with as few as one training sample.

Pith tools