Pith. sign in

REVIEW 13 cited by

Inpaint Anything: Segment Anything Meets Image Inpainting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.06790 v1 pith:7PF76WUY submitted 2023-04-13 cs.CV

classification cs.CV
keywords anythingimageinpaintinpaintingusersfillfillinghole
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern image inpainting systems, despite the significant progress, often struggle with mask selection and holes filling. Based on Segment-Anything Model (SAM), we make the first attempt to the mask-free image inpainting and propose a new paradigm of ``clicking and filling'', which is named as Inpaint Anything (IA). The core idea behind IA is to combine the strengths of different models in order to build a very powerful and user-friendly pipeline for solving inpainting-related problems. IA supports three main features: (i) Remove Anything: users could click on an object and IA will remove it and smooth the ``hole'' with the context; (ii) Fill Anything: after certain objects removal, users could provide text-based prompts to IA, and then it will fill the hole with the corresponding generative content via driving AIGC models like Stable Diffusion; (iii) Replace Anything: with IA, users have another option to retain the click-selected object and replace the remaining background with the newly generated scenes. We are also very willing to help everyone share and promote new projects based on our Inpaint Anything (IA). Our codes are available at https://github.com/geekyutao/Inpaint-Anything.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A SAM-based model with shape/intensity adapters, neighboring feature aggregation, and wavelet detail enhancement achieves 73.6 mAP on a new 47-species zooplankton microscopy dataset.

  2. Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A 3D-Gaussian-plus-diffusion pipeline translates multi-modal glove HOI videos into photorealistic bare-hand videos, yielding the HandSense dataset that improves contact estimation and occluded tracking.

  3. 2D Gaussian Splatting with Semantic Alignment for Image Inpainting

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 2D Gaussian Splatting encoder-rasterization network with DINO-based semantic alignment achieves competitive image inpainting results.

  4. Neural Scene Designer: Self-Styled Semantic Image Manipulation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    NSD uses a contrastively learned style embedding from the input image itself, fed through a second cross-attention branch, to make diffusion-based inpainting results match the surrounding scene's style.

  5. MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips

    cs.CV 2025-08 conditional novelty 6.0 of 10

    MagicHOI integrates a novel view synthesis diffusion prior with a visibility-aware weighting strategy to reconstruct accurate hand-object 3D shapes from short monocular videos with partial object visibility.

  6. SceneLoom: Communicating Data with Scene Context

    cs.HC 2025-07 conditional novelty 6.0 of 10

    SceneLoom guides a vision-language model through a design space derived from 54 data videos to generate chart-in-image designs aligned with user narrative intent.

  7. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  8. ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A few-shot segmentation framework that injects reference-image prototypes into SAM's decoder and image encoder, eliminating per-image manual prompts and improving remote sensing segmentation accuracy.

  9. GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A generative model reconstructs complete, occlusion-free 3D objects from sparse unposed views by combining 2D amodal inpainting with stereo-point-cloud-conditioned cross-attention.

  10. EfficientIML: Efficient High-Resolution Image Manipulation Localization

    cs.CV 2025-09 conditional novelty 5.0 of 10

    EfficientIML combines a lightweight RWKV-based backbone with multi-scale supervision to accurately localize high-resolution diffusion-based image forgeries, while introducing the SIF dataset for training and evaluation.

  11. GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design

    cs.HC 2025-08 conditional novelty 5.0 of 10

    GenTune improves AI image refinement by tracing image regions back to prompt labels and allowing element-level, semantic-guided edits.

  12. Preserve Anything: Controllable Image Synthesis with Object Preservation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Preserve Anything introduces N-channel conditioning to ControlNet, combining object masks, background edge layouts, and lighting gradients, and reports improved FID and user-study scores for object-preserving image synthesis.

  13. PairEdit: Learning Semantic Variations for Exemplar-based Image Editing

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PairEdit trains two LoRA adapters on a pretrained diffusion model to capture the semantic direction between paired source-target images, enabling text-free, controllable image editing from as few as one pair.

Pith tools