Pith. sign in

REVIEW 6 cited by

MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08000 v3 pith:TVF3YZJB submitted 2024-08-15 cs.CV

classification cs.CV
keywords mvinpaintermulti-viewsynthesisattentioncameraeditingguidancein-the-wild
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Novel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which are discouraged from generalizing to challenging in-the-wild scenes and fail to be employed with 2D synthesis directly. Moreover, these methods heavily depended on camera poses, limiting their real-world applications. To overcome these issues, we propose MVInpainter, re-formulating the 3D editing as a multi-view 2D inpainting task. Specifically, MVInpainter partially inpaints multi-view images with the reference guidance rather than intractably generating an entirely novel view from scratch, which largely simplifies the difficulty of in-the-wild NVS and leverages unmasked clues instead of explicit pose conditions. To ensure cross-view consistency, MVInpainter is enhanced by video priors from motion components and appearance guidance from concatenated reference key&value attention. Furthermore, MVInpainter incorporates slot attention to aggregate high-level optical flow features from unmasked regions to control the camera movement with pose-free training and inference. Sufficient scene-level experiments on both object-centric and forward-facing datasets verify the effectiveness of MVInpainter, including diverse tasks, such as multi-view object removal, synthesis, insertion, and replacement. The project page is https://ewrfcas.github.io/MVInpainter/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic-Guided Progressive Object Removal with Gaussian Splatting

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Semantic block matching via DINOv2 plus selective high-frequency refinement yields higher-fidelity, multi-view-consistent object removal inside 3D Gaussian Splatting than prior one-shot Gaussian or NeRF inpainters.

  2. SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models

    cs.RO 2025-11 conditional novelty 6.0 of 10

    SplatSearch combines sparse-view 3D Gaussian Splatting, multi-view diffusion inpainting, and semantic/visual frontier scoring to achieve viewpoint-invariant instance image-goal navigation in unknown environments.

  3. DSG-World: Learning a 3D Gaussian World Model from Dual State Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DSG-World builds two segmented 3D Gaussian fields from two scene states and trains them with mutual consistency, enabling novel-state simulation without inpainting or dense capture.

  4. SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A depth-guided Gaussian Splatting inpainting method with soft depth clustering and selective guided refinement achieves modest quality gains and 24.5% faster training over GScream on SPIn-NeRF.

  5. DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D Inpainting

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DiGA3D performs text-guided 3D inpainting (removal, re-texturing, replacement) with a coarse-to-fine diffusion propagation scheme to improve multi-view appearance and geometry consistency.

  6. VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal

    cs.GR 2025-06 conditional novelty 5.0 of 10

    VEIGAR is a pipeline for 3D object removal in Gaussian Splatting that uses deep stereo depth projection and a scale-invariant depth loss to achieve faster training and comparable quality to prior state-of-the-art.

Pith tools