Pith. sign in

REVIEW 12 cited by

Large Scale Image Completion via Co-Modulated Generative Adversarial Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.10428 v1 pith:ED3GLWLS submitted 2021-03-18 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords completionimagegenerativeadversarialconditionalimagesnetworkspropose
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Numerous task-specific variants of conditional generative adversarial networks have been developed for image completion. Yet, a serious limitation remains that all existing algorithms tend to fail when handling large-scale missing regions. To overcome this challenge, we propose a generic new approach that bridges the gap between image-conditional and recent modulated unconditional generative architectures via co-modulation of both conditional and stochastic style representations. Also, due to the lack of good quantitative metrics for image completion, we propose the new Paired/Unpaired Inception Discriminative Score (P-IDS/U-IDS), which robustly measures the perceptual fidelity of inpainted images compared to real images via linear separability in a feature space. Experiments demonstrate superior performance in terms of both quality and diversity over state-of-the-art methods in free-form image completion and easy generalization to image-to-image translation. Code is available at https://github.com/zsyzzsoft/co-mod-gan.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal

    cs.CV 2025-12 conditional novelty 6.0 of 10

    EraseLoRA removes masked objects by having an MLLM separate target, non-target foreground, and background, then test-time LoRA optimization aggregates background subtypes to reconstruct the occluded region.

  2. Neural Scene Designer: Self-Styled Semantic Image Manipulation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    NSD uses a contrastively learned style embedding from the input image itself, fed through a second cross-attention branch, to make diffusion-based inpainting results match the surrounding scene's style.

  3. DreamPainter: Image Background Inpainting for E-commerce Scenarios

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DreamPainter introduces a two-stage diffusion framework trained on a new synthetic e-commerce dataset, DreamEcom-400K, that outperforms open-source inpainting baselines on background generation with text and reference...

  4. OutDreamer: Video Outpainting with a Diffusion Transformer

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OutDreamer couples a diffusion transformer with mask-driven self-attention and a latent alignment loss to outpaint videos in a zero-shot manner, exceeding prior zero-shot baselines on standard benchmarks.

  5. Towards Seamless Borders: A Method for Mitigating Inconsistencies in Image Inpainting and Outpainting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-step training loss plus a fine-tuned VAE reduces color and structure discontinuities at mask boundaries in diffusion image inpainting and outpainting.

  6. HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation

    cs.GR 2025-04 conditional novelty 6.0 of 10

    HiScene generates compositional 3D scenes by treating a room as an object under isometric view, then decomposing and regenerating each instance with video-diffusion amodal completion.

  7. BrushEdit: All-In-One Image Inpainting and Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.

  8. 3D-Consistent Image Inpainting with Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion inpainting model conditioned on a second viewpoint of the same scene produces 3D-consistent fills for occluded regions without 3D supervision.

  9. PainterNet: Adaptive Image Inpainting with Actual-Token Attention and Diverse Mask Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PainterNet is a diffusion-model plugin that uses local prompts, attention supervision, and diverse masks to improve text-consistent image inpainting.

  10. Gimbal360: Canonicalizing Planar Diffusion for Spherical Panorama Completion

    cs.CV 2026-03 reject novelty 5.0 of 10

    Gimbal360 completes 360° panoramas from unposed perspective images by rigidly auto-leveling inputs and training diffusion with a Siamese shift-equivariance loss to preserve ERP seam continuity.

  11. SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A depth-guided Gaussian Splatting inpainting method with soft depth clustering and selective guided refinement achieves modest quality gains and 24.5% faster training over GScream on SPIn-NeRF.

  12. MagicRoad: Semantic-Aware 3D Road Surface Reconstruction via Obstacle Inpainting

    cs.CV 2025-07 reject novelty 5.0 of 10

    MagicRoad combines video inpainting, semantic color harmonization, and 2D Gaussian surfels to reconstruct clean bird's eye view road surfaces from driving video.

Pith tools