Pith. sign in

REVIEW 17 cited by

OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.07397 v3 pith:ZQ3BQ32Y submitted 2025-01-13 cs.CV

classification cs.CV
keywords imagesobjecteffectsobjectsomnieraserartifactscontentdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inpainting algorithms have achieved remarkable progress in removing objects from images, yet still face two challenges: 1) struggle to handle the object's visual effects such as shadow and reflection; 2) easily generate shape-like artifacts and unintended content. In this paper, we propose Video4Removal, a large-scale dataset comprising over 100,000 high-quality samples with realistic object shadows and reflections. By constructing object-background pairs from video frames with off-the-shelf vision models, the labor costs of data acquisition can be significantly reduced. To avoid generating shape-like artifacts and unintended content, we propose Object-Background Guidance, an elaborated paradigm that takes both the foreground object and background images. It can guide the diffusion process to harness richer contextual information. Based on the above two designs, we present OmniEraser, a novel method that seamlessly removes objects and their visual effects using only object masks as input. Extensive experiments show that OmniEraser significantly outperforms previous methods, particularly in complex in-the-wild scenes. And it also exhibits a strong generalization ability in anime-style images. Datasets, models, and codes will be published.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Impostor: An Agent-Curated Benchmark for Realistic AIGC Manipulation Localization

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Introduces the Impostor benchmark dataset for localizing AIGC image manipulations via agent curation and the PANet model that uses phase and semantic consistency for better detection.

  2. PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    PROVE proposes RC metrics for perceptual removal coherence and releases PROVE-Bench to better align automatic scores with human judgments on object removal tasks.

  3. PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media

    cs.CV 2026-05 conditional novelty 7.0 of 10

    Removal Coherence (RC) metrics, which compare local feature distributions in masked versus background regions via sliding-window MMD, align with human judgments of object-removal quality substantially better than exis...

  4. SceneForge: Structured World Supervision from 3D Interventions

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    SceneForge creates intervention-consistent multimodal supervision from editable 3D world states, yielding improved object and scene removal performance on benchmarks.

  5. A Unified and Controllable Framework for Layered Image Generation with Visual Effects

    cs.CV 2026-01 unverdicted novelty 7.0 of 10

    LASAGNA produces layered images with integrated visual effects in a single pass, enabling drift-free edits via alpha compositing while releasing a 48K dataset and a 242-sample benchmark.

  6. DORS: Dynamic Attention Routing for Diffusion-based Object Removal in Dense Scenes

    cs.CV 2026-07 conditional novelty 6.0 of 10

    DORS edits self-attention during diffusion denoising so masked pixels ignore similar surrounding instances, reporting large artifact reductions in dense-scene object removal.

  7. OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    OSOR is a one-step diffusion inpainting method using an occupancy-guided discriminator, alpha head, and semantic-anchored verification pipeline to achieve effect-aware object removal, outperforming multi-step baseline...

  8. Concept Removal for Frontier Image Generative Models

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    A transcoder-based in-place replacement of the bottleneck layer enables selective concept removal in modern diffusion and autoregressive image models without degrading output quality.

  9. FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    FlashClear delivers up to 122x faster object removal than prior diffusion models via adversarial step distillation and asymmetric attention caching while preserving visual quality.

  10. FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    FlashClear achieves up to 8.26x speedup over its base diffusion model and 122x over OmniPaint for image object removal via region-aware adversarial distillation and foreground-prioritized caching while claiming to mai...

  11. Relit-LiVE: Relight Video by Jointly Learning Environment Video

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Relit-LiVE jointly predicts relit videos and viewpoint-aligned environment maps inside a single diffusion process to achieve physically consistent video relighting without camera pose input.

  12. GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    GRAFT amortizes human-scene fitting into a recurrent transformer that predicts interaction gradients via body-anchored geometric probes, delivering optimization-level interaction quality at 50x lower runtime.

  13. EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal

    cs.CV 2025-12 conditional novelty 6.0 of 10

    EraseLoRA removes masked objects by having an MLLM separate target, non-target foreground, and background, then test-time LoRA optimization aggregates background subtypes to reconstruct the occluded region.

  14. Mask Consistency Regularization in Object Removal

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A mask-consistency training loss, enforcing equal predictions across dilated and reshaped masks, is proposed to reduce hallucination and mask-shape bias in diffusion-based object removal.

  15. ROSE: Remove Objects with Side Effects in Videos

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video inpainting model trained on 3D-rendered pairs removes objects together with their shadows, reflections, and other side effects, plus a new benchmark.

  16. GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    GenEraser proposes MC-MoE with bipartite text guidance, LD-CFG fusion, and a decoupled locator-preserver architecture for generalizable video object and effect removal, claiming 2.16 dB and 1.44 dB gains on ROSE and V...

  17. GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

    cs.CV 2026-04 conditional novelty 4.0 of 10

    Correcting the elasticity choices and decomposition method of a prior study shows the US embargo can explain a substantial share of Cuba's post-1959 economic underperformance.

Pith tools