Pith. sign in

REVIEW 6 cited by

MVDiffusion++: A Dense High-resolution Multi-view Diffusion Model for Single or Sparse-view 3D Object Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12712 v3 pith:RZUY4OJ5 submitted 2024-02-20 cs.CV

classification cs.CV
keywords mvdiffusiondensehigh-resolutionobjectreconstructionviewviewsarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a neural architecture MVDiffusion++ for 3D object reconstruction that synthesizes dense and high-resolution views of an object given one or a few images without camera poses. MVDiffusion++ achieves superior flexibility and scalability with two surprisingly simple ideas: 1) A ``pose-free architecture'' where standard self-attention among 2D latent features learns 3D consistency across an arbitrary number of conditional and generation views without explicitly using camera pose information; and 2) A ``view dropout strategy'' that discards a substantial number of output views during training, which reduces the training-time memory footprint and enables dense and high-resolution view synthesis at test time. We use the Objaverse for training and the Google Scanned Objects for evaluation with standard novel view synthesis and 3D reconstruction metrics, where MVDiffusion++ significantly outperforms the current state of the arts. We also demonstrate a text-to-3D application example by combining MVDiffusion++ with a text-to-image generative model. The project page is at https://mvdiffusion-plusplus.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A multi-view diffusion pipeline that segments 3D objects into parts, completes occluded or invisible parts, and reconstructs them into a compositional 3D asset.

  2. LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LiftImage3D generates small-motion video clips from one image, registers them with MASt3R, and fits a distortion-aware 3D Gaussian field whose canonical scene renders new views.

  3. CheapNVS: Real-Time On-Device Narrow-Baseline Novel View Synthesis

    cs.CV 2025-01 conditional novelty 5.0 of 10

    CheapNVS performs narrow-baseline single-view novel view synthesis on mobile devices by learning warping and inpainting in parallel from a shared latent space.

  4. UnCommon Objects in 3D

    cs.CV 2025-01 conditional novelty 5.0 of 10

    uCO3D is a large, diverse, high-quality real-object video dataset with 3D annotations that improves training of feedforward 3D reconstruction and text-to-3D models.

  5. Pragmatist: Multiview Conditional Diffusion Models for High-Fidelity 3D Reconstruction from Unposed Sparse Views

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Pragmatist turns sparse unposed photos of an object into a high-fidelity 3D mesh by generating consistent canonical views with a diffusion model, reconstructing a triplane mesh, then refining camera poses and texture ...

  6. Make-A-Texture: Fast Shape-Aware Texture Generation in 3 Seconds

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A texture-generation pipeline that produces 1024x1024 textures from text in 3.07 seconds on an H100, with quality comparable to SyncMVD and other prior methods.

Pith tools