REVIEW 13 cited by
Inpaint Anything: Segment Anything Meets Image Inpainting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Modern image inpainting systems, despite the significant progress, often struggle with mask selection and holes filling. Based on Segment-Anything Model (SAM), we make the first attempt to the mask-free image inpainting and propose a new paradigm of ``clicking and filling'', which is named as Inpaint Anything (IA). The core idea behind IA is to combine the strengths of different models in order to build a very powerful and user-friendly pipeline for solving inpainting-related problems. IA supports three main features: (i) Remove Anything: users could click on an object and IA will remove it and smooth the ``hole'' with the context; (ii) Fill Anything: after certain objects removal, users could provide text-based prompts to IA, and then it will fill the hole with the corresponding generative content via driving AIGC models like Stable Diffusion; (iii) Replace Anything: with IA, users have another option to retain the click-selected object and replace the remaining background with the newly generated scenes. We are also very willing to help everyone share and promote new projects based on our Inpaint Anything (IA). Our codes are available at https://github.com/geekyutao/Inpaint-Anything.
Forward citations
Cited by 13 Pith papers
-
ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
A SAM-based model with shape/intensity adapters, neighboring feature aggregation, and wavelet detail enhancement achieves 73.6 mAP on a new 47-species zooplankton microscopy dataset.
-
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
A 3D-Gaussian-plus-diffusion pipeline translates multi-modal glove HOI videos into photorealistic bare-hand videos, yielding the HandSense dataset that improves contact estimation and occluded tracking.
-
2D Gaussian Splatting with Semantic Alignment for Image Inpainting
A 2D Gaussian Splatting encoder-rasterization network with DINO-based semantic alignment achieves competitive image inpainting results.
-
Neural Scene Designer: Self-Styled Semantic Image Manipulation
NSD uses a contrastively learned style embedding from the input image itself, fed through a second cross-attention branch, to make diffusion-based inpainting results match the surrounding scene's style.
-
MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips
MagicHOI integrates a novel view synthesis diffusion prior with a visibility-aware weighting strategy to reconstruct accurate hand-object 3D shapes from short monocular videos with partial object visibility.
-
SceneLoom: Communicating Data with Scene Context
SceneLoom guides a vision-language model through a design space derived from 54 data videos to generate chart-in-image designs aligned with user narrative intent.
-
RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.
-
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
A few-shot segmentation framework that injects reference-image prototypes into SAM's decoder and image encoder, eliminating per-image manual prompts and improving remote sensing segmentation accuracy.
-
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
A generative model reconstructs complete, occlusion-free 3D objects from sparse unposed views by combining 2D amodal inpainting with stereo-point-cloud-conditioned cross-attention.
-
EfficientIML: Efficient High-Resolution Image Manipulation Localization
EfficientIML combines a lightweight RWKV-based backbone with multi-scale supervision to accurately localize high-resolution diffusion-based image forgeries, while introducing the SIF dataset for training and evaluation.
-
GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
GenTune improves AI image refinement by tracing image regions back to prompt labels and allowing element-level, semantic-guided edits.
-
Preserve Anything: Controllable Image Synthesis with Object Preservation
Preserve Anything introduces N-channel conditioning to ControlNet, combining object masks, background edge layouts, and lighting gradients, and reports improved FID and user-study scores for object-preserving image synthesis.
-
PairEdit: Learning Semantic Variations for Exemplar-based Image Editing
PairEdit trains two LoRA adapters on a pretrained diffusion model to capture the semantic direction between paired source-target images, enabling text-free, controllable image editing from as few as one pair.
Discussion (0). Sign in to comment.