Pith. sign in

REVIEW 31 cited by

Resolution-robust Large Mask Inpainting with Fourier Convolutions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.07161 v2 pith:KQAM3OCX submitted 2021-09-15 cs.CV eess.IV

classification cs.CVeess.IV
keywords inpaintinglargefieldlamanetworkreceptiveachievesconvolutions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern image inpainting systems, despite the significant progress, often struggle with large missing areas, complex geometric structures, and high-resolution images. We find that one of the main reasons for that is the lack of an effective receptive field in both the inpainting network and the loss function. To alleviate this issue, we propose a new method called large mask inpainting (LaMa). LaMa is based on i) a new inpainting network architecture that uses fast Fourier convolutions (FFCs), which have the image-wide receptive field; ii) a high receptive field perceptual loss; iii) large training masks, which unlocks the potential of the first two components. Our inpainting network improves the state-of-the-art across a range of datasets and achieves excellent performance even in challenging scenarios, e.g. completion of periodic structures. Our model generalizes surprisingly well to resolutions that are higher than those seen at train time, and achieves this at lower parameter&time costs than the competitive baselines. The code is available at \url{https://github.com/saic-mdal/lama}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AsyncPatch Diffusion: spatially-flexible image generation

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    AsyncPatch Diffusion introduces asynchronous per-region noise levels in diffusion models, proves a valid ELBO, and uses a controlled sampler to support spatially adaptive generation and native inpainting.

  2. FantasyID: A dataset for detecting digital manipulations of ID-documents

    cs.CV 2025-07 conditional novelty 7.0 of 10

    The new FantasyID benchmark shows that current forgery detectors miss roughly half of face-swapped ID cards at a 10% false positive rate, while performing better on text edits.

  3. FocusView: Understanding and Customizing Informational Video Watching Experiences for Viewers with ADHD

    cs.HC 2025-07 conditional novelty 7.0 of 10

    FocusView, a customizable video interface, improved self-reported viewability of informational videos for 12 participants with ADHD.

  4. High-Resolution Image Synthesis with Latent Diffusion Models

    cs.CV 2021-12 conditional novelty 7.0 of 10

    Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrai...

  5. Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An FFC-based GAN guided by SAR and topographic data reconstructs 30 m clear-sky land-surface temperature in cloud-covered Landsat pixels with typical RMSE 0.8–1.8 K.

  6. LENS: LLM-guided Environment Simplification for Planning and Control in Clutter

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A vision-language-model-based prune-and-merge abstraction improves success and runtime for TAMP, contact-implicit MPC, and a VLA policy in cluttered tabletop manipulation.

  7. What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Erasing objects from front-camera images shows Alpamayo 1's trajectories depend most on large vehicles, pedestrians, and traffic lights, but attributions are seed-unstable and some effects reach the output without tou...

  8. Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Moebius introduces a compressed diffusion inpainting model using Local-λ Mix Interaction blocks and latent-space multi-granularity distillation to reach 10B-level quality with 0.22B parameters.

  9. GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    GS-Playground delivers a high-throughput photorealistic simulator for vision-informed robot learning via parallel physics integrated with batch 3D Gaussian Splatting at 10^4 FPS and an automated Real2Sim workflow for ...

  10. Multimodal Language Models Cannot Spot Spatial Inconsistencies

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Multimodal LLMs significantly underperform humans at spotting objects that break 3D consistency in multi-view image pairs.

  11. Generative deep learning improves reconstruction of global historical climate records

    physics.geo-ph 2026-02 unverdicted novelty 6.0 of 10

    A probabilistic generative deep learning framework reconstructs global historical climate fields from 1850 onward, revealing higher early 20th-century warming driven by stronger polar trends and localized modern hotsp...

  12. ROSE: Remove Objects with Side Effects in Videos

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video inpainting model trained on 3D-rendered pairs removes objects together with their shadows, reflections, and other side effects, plus a new benchmark.

  13. DreamPainter: Image Background Inpainting for E-commerce Scenarios

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DreamPainter introduces a two-stage diffusion framework trained on a new synthetic e-commerce dataset, DreamEcom-400K, that outperforms open-source inpainting baselines on background generation with text and reference...

  14. Scalable and Realistic Virtual Try-on Application for Foundation Makeup with Kubelka-Munk Theory

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A virtual try-on framework approximates Kubelka-Munk optical blending with a Taylor expansion, enabling realistic foundation previews using only e-commerce product data.

  15. Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

    cs.CV 2024-01 unverdicted novelty 6.0 of 10

    Grounded SAM integrates Grounding DINO and SAM to support text-prompted open-world detection and segmentation, achieving 48.7 mean AP on SegInW zero-shot with the base detector and huge segmenter.

  16. Scaling Robot Learning with Semantically Imagined Experience

    cs.RO 2023-02 unverdicted novelty 6.0 of 10

    Augmenting robot datasets via diffusion-based semantic inpainting enables manipulation policies to solve unseen tasks with new objects and improves robustness to novel distractors.

  17. 3D-GIMP: When 3D Gaussian Inpainting Meets PatchMatch

    cs.CV 2026-07 conditional novelty 5.0 of 10

    3D-GIMP removes objects from 3D Gaussian Splatting scenes by inpainting one reference view and propagating it to all views via a 3D-aware PatchMatch field, cutting optimization time from ~1 hour to ~6 minutes.

  18. Restore3D: Breathing Life into Broken Objects with Shape and Texture Restoration

    cs.CV 2026-07 unverdicted novelty 5.0 of 10

    Restore3D restores shape and texture of broken 3D objects via multi-view image refinement with a Mask Self-Perceiver and coarse-to-fine mesh reconstruction, outperforming baselines on synthetic and real benchmarks.

  19. GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    GASE automates high-fidelity simulation scene reconstruction from multi-view panoramic videos via Gaussian splatting, object extraction, and inpainting, yielding robot policies with under 10% performance gap versus re...

  20. Benchmarking Single-Step Inpainting Methods for Multi-Object 3D Gaussian Splatting Scenes

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Reconstruction-based 2D inpainters outperform generative ones for 3D consistency in 3DGS object removal; scratch initialization beats finetuning; supported by a new multi-object dataset with ground truth and occlusions.

  21. PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

    cs.GR 2026-05 unverdicted novelty 5.0 of 10

    TelePhysics is a training-free pipeline that builds a unified 3D scene model from one photo and then runs decoupled physics simulation to produce controllable, penetration-free multi-object videos.

  22. PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

    cs.GR 2026-05 conditional novelty 5.0 of 10

    A training-free pipeline reconstructs a single image into a physically simulated 3D scene and re-renders simulated frames into a controllable, photorealistic video.

  23. DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    DecoRec decomposes single-view 3D scene reconstruction into per-object diffusion reconstructions followed by a differentiable rendering and diffusion-guided merging pipeline.

  24. STEMTOX: From Collaborative Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A new dataset and an entropy-guided multi-task model show that collaborative tags improve fine-grained toxic meme detection.

  25. SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

    cs.CV 2023-11 unverdicted novelty 5.0 of 10

    SPHINX improves multi-modal LLMs through joint mixing of weights, tasks, and visual embeddings from varied sources to achieve stronger alignment and multi-purpose capabilities.

  26. CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

    cs.CV 2026-05 reject novelty 4.0 of 10

    The paper's stated CA-World counterfactual claim is absent from the body, which instead describes the SAM3D-Phys pipeline for multi-object interactive reconstruction and simulation.

  27. CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    SAM3D-Phys recovers complete simulatable object geometries from incomplete real-world scene reconstructions by combining SAM3D generative priors with physics-constrained spatial optimization and mask-guided appearance...

  28. Neural Field Representations of Mobile Computational Photography

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Fitting neural fields directly to raw phone bursts reconstructs depth, separates reflections and occluders, and stitches panoramas, outperforming the compared baselines on the thesis's benchmarks.

  29. PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback

    cs.CV 2025-06 unverdicted novelty 4.0 of 10

    PrefPaint uses D3PO and a Model Tree web interface to incorporate gastroenterologist feedback into Stable Diffusion inpainting, producing anatomically accurate polyp images that outperform prior methods in user studies.

  30. WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

    cs.SD 2025-08 conditional novelty 3.0 of 10

    WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.

  31. Enhancing non-Rigid 3D Model Deformations Using Mesh-based Gaussian Splatting

    cs.GR 2025-07 reject novelty 2.0 of 10

    A proposal to combine 3D Gaussian splatting, SAM segmentation, GS2Mesh conversion, LLM-based material assignment, and XPBD physics into a mesh-based editing pipeline, with no experimental validation.

Pith tools